{
  "id": 511535,
  "title": "5th solution",
  "url": "/competitions/birdclef-2024/discussion/511535",
  "author_name": "coolz",
  "post_date": "2024-06-11T05:51:40.109000",
  "votes": 57,
  "comment_count": 39,
  "views": 0,
  "content": "<p>First of all, thanks to the organizers for hosting this competition.<br>\nAnd congrats to all winners.</p>\n<h3>Data</h3>\n<p>2024 data</p>\n<h3>Model</h3>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>backbone</th>\n<th>weights</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raw signale</td>\n<td>hgnetb0</td>\n<td>3/5 folds</td>\n<td>0.720354</td>\n<td>0.666671</td>\n</tr>\n<tr>\n<td>spectrum</td>\n<td>efficientnetb0</td>\n<td>1/5 folds</td>\n<td>0.708198</td>\n<td>0.672360</td>\n</tr>\n<tr>\n<td>raw+spectrum</td>\n<td>hgnetb0+effb0</td>\n<td>1/5 folds</td>\n<td>0.698408</td>\n<td>0.660503</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>…</td>\n<td>0.5+0.4+0.1</td>\n<td>0.743960</td>\n<td>0.687173</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>1. raw signal model</li>\n</ul>\n<pre><code>\n\n = x.view(bs, -, )           \n\n = torch.transpose(x, , )    \n\n = x.view(bs, , -, )         \n = self.backbone(x)   \n</code></pre>\n<ul>\n<li>2. spectrum model</li>\n</ul>\n<pre><code>torchaudio.transforms.MelSpectrogram(\n              32000,\n              =512,\n              =0,\n              =16000,\n              =2048*2,\n              =512,\n              =,\n          ),\ntorchaudio.transforms.AmplitudeToDB(=80)\n</code></pre>\n<ul>\n<li>3. mix model</li>\n</ul>\n<pre><code>  = self.raw_model(x)\n = self.spec_model(x)\n = torch.cat([raw_f,spec_f],dim=)\n=self.fc(feature)\n\n\n</code></pre>\n<h3>Preprocess</h3>\n<ul>\n<li>1. wave/max(wave) if max(wave)&gt;1</li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>XY cut out for spectrum model</li>\n</ul>\n<h3>Train</h3>\n<ul>\n<li>1. Random sample 5 seconds data for trainning, <br>\nfirst 5 seconds data for validation.</li>\n<li>2. Mixup p=1.</li>\n<li>3. BCEWithLogitsLoss</li>\n</ul>\n<h3>inference</h3>\n<ul>\n<li>openvino</li>\n</ul>\n<h3>extra</h3>\n<p>I make a post to discussion the method feeding raw signal to vision model. Here is the link <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511763\" target=\"_blank\">good luck</a><br>\ninference <a href=\"https://www.kaggle.com/code/cooolz/5nd-solution\" target=\"_blank\">notebook</a></p>",
  "messages": [
    {
      "id": 2866034,
      "postDate": "2024-06-11T05:51:40.110Z",
      "content": "<p>First of all, thanks to the organizers for hosting this competition.<br>\nAnd congrats to all winners.</p>\n<h3>Data</h3>\n<p>2024 data</p>\n<h3>Model</h3>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>backbone</th>\n<th>weights</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raw signale</td>\n<td>hgnetb0</td>\n<td>3/5 folds</td>\n<td>0.720354</td>\n<td>0.666671</td>\n</tr>\n<tr>\n<td>spectrum</td>\n<td>efficientnetb0</td>\n<td>1/5 folds</td>\n<td>0.708198</td>\n<td>0.672360</td>\n</tr>\n<tr>\n<td>raw+spectrum</td>\n<td>hgnetb0+effb0</td>\n<td>1/5 folds</td>\n<td>0.698408</td>\n<td>0.660503</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>…</td>\n<td>0.5+0.4+0.1</td>\n<td>0.743960</td>\n<td>0.687173</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>1. raw signal model</li>\n</ul>\n<pre><code>\n\n = x.view(bs, -, )           \n\n = torch.transpose(x, , )    \n\n = x.view(bs, , -, )         \n = self.backbone(x)   \n</code></pre>\n<ul>\n<li>2. spectrum model</li>\n</ul>\n<pre><code>torchaudio.transforms.MelSpectrogram(\n              32000,\n              =512,\n              =0,\n              =16000,\n              =2048*2,\n              =512,\n              =,\n          ),\ntorchaudio.transforms.AmplitudeToDB(=80)\n</code></pre>\n<ul>\n<li>3. mix model</li>\n</ul>\n<pre><code>  = self.raw_model(x)\n = self.spec_model(x)\n = torch.cat([raw_f,spec_f],dim=)\n=self.fc(feature)\n\n\n</code></pre>\n<h3>Preprocess</h3>\n<ul>\n<li>1. wave/max(wave) if max(wave)&gt;1</li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>XY cut out for spectrum model</li>\n</ul>\n<h3>Train</h3>\n<ul>\n<li>1. Random sample 5 seconds data for trainning, <br>\nfirst 5 seconds data for validation.</li>\n<li>2. Mixup p=1.</li>\n<li>3. BCEWithLogitsLoss</li>\n</ul>\n<h3>inference</h3>\n<ul>\n<li>openvino</li>\n</ul>\n<h3>extra</h3>\n<p>I make a post to discussion the method feeding raw signal to vision model. Here is the link <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/511763\" target=\"_blank\">good luck</a><br>\ninference <a href=\"https://www.kaggle.com/code/cooolz/5nd-solution\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "\n\nFirst of all, thanks to the organizers for hosting this competition.\nAnd congrats to all winners.\n\n\n### Data\n2024 data\n\n### Model\n| model                          | backbone       | weights     | public   | private  |\n|--------------------------------|----------------|-------------|----------|----------|\n| raw signale                    | hgnetb0        | 3/5 folds   | 0.720354 | 0.666671 |\n| spectrum                       | efficientnetb0 | 1/5 folds   | 0.708198    | 0.672360    |\n| raw+spectrum                   | hgnetb0+effb0  | 1/5 folds   | 0.698408 | 0.660503 |\n| ensemble                       | ...            | 0.5+0.4+0.1 | 0.743960 | 0.687173 |\n\n+ 1. raw signal model\n```\n# x bsx160000\n\n#bsx80000x2\nx = x.view(bs, -1, 2)           \n#bsx2x80000\nx = torch.transpose(x, 2, 1)    \n#bsx2x1000x80\nx = x.view(bs, 2, -1, 80)         \nfeature = self.backbone(x)   \n```\n+ 2. spectrum model\n```\ntorchaudio.transforms.MelSpectrogram(\n                32000,\n                n_mels=512,\n                f_min=0,\n                f_max=16000,\n                n_fft=2048*2,\n                hop_length=512,\n                normalized=True,\n            ),\n\ntorchaudio.transforms.AmplitudeToDB(top_db=80)\n```\n+ 3. mix model\n\n```commandline\nraw_f  = self.raw_model(x)\nspec_f = self.spec_model(x)\nfeature = torch.cat([raw_f,spec_f],dim=1)\nx=self.fc(feature)\n\n## spectrum is resized to 256x256 for fast speed, \n## other params for mel spectrum is the same with spectrm model.\n```\n### Preprocess\n+ 1. wave/max(wave) if max(wave)>1\n### Augmentation\n+ XY cut out for spectrum model\n### Train\n+ 1. Random sample 5 seconds data for trainning, \nfirst 5 seconds data for validation.\n+ 2. Mixup p=1.\n+ 3. BCEWithLogitsLoss\n\n### inference\n+ openvino\n\n\n### extra\nI make a post to discussion the method feeding raw signal to vision model. Here is the link [good luck](https://www.kaggle.com/competitions/birdclef-2024/discussion/511763)\n\n\ninference [notebook](https://www.kaggle.com/code/cooolz/5nd-solution)\n",
      "votes": 57
    },
    {
      "id": 2866062,
      "postDate": "2024-06-11T06:04:18.460Z",
      "content": "<p>Regarding the \"raw model\", how did you deal with it to achieve good results? I rarely see people using raw signals in Bird competitions.</p>",
      "rawMarkdown": "Regarding the \"raw model\", how did you deal with it to achieve good results? I rarely see people using raw signals in Bird competitions.",
      "votes": 3,
      "replies": [
        {
          "id": 2866242,
          "postDate": "2024-06-11T08:08:36.480Z",
          "content": "<p>I add some comments one the pse codes, i think you can get it now.</p>",
          "rawMarkdown": "I add some comments one the pse codes, i think you can get it now.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2869294,
      "postDate": "2024-06-13T02:25:50.187Z",
      "content": "<p>Good work. I was betting you'd use packed raw waveform as you did in HMS competition!</p>\n<p>Congrats on becoming a GM. It was a matter of time given what you did in HMS.</p>\n<p>Did you compare openvino accuracy with pytorch models?  Asking because we found that it loses some accuracy and we didn't use it as a result.</p>",
      "rawMarkdown": "Good work. I was betting you'd use packed raw waveform as you did in HMS competition!\n\nCongrats on becoming a GM. It was a matter of time given what you did in HMS.\n\nDid you compare openvino accuracy with pytorch models?  Asking because we found that it loses some accuracy and we didn't use it as a result.",
      "votes": 1,
      "replies": [
        {
          "id": 2869391,
          "postDate": "2024-06-13T04:32:40.203Z",
          "content": "<p>Thanks.</p>\n<p>I noticed that you'd mentioned the precision loss with openvino. <br>\nI manully check some of the outputs of pytorch, onnxruntime, and openvino, and notice very small differences, but didn't try to sub with onnxruntime. By the way, as the the score is rank relavent, i use logits as final output. </p>",
          "rawMarkdown": "Thanks.\n\nI noticed that you'd mentioned the precision loss with openvino. \nI manully check some of the outputs of pytorch, onnxruntime, and openvino, and notice very small differences, but didn't try to sub with onnxruntime. By the way, as the the score is rank relavent, i use logits as final output. ",
          "votes": 1,
          "replies": [
            {
              "id": 2870282,
              "postDate": "2024-06-13T13:48:26.987Z",
              "content": "<p>I used logits too until we found that smoothing worked better when applied to probabilities.</p>",
              "rawMarkdown": "I used logits too until we found that smoothing worked better when applied to probabilities.",
              "votes": 1
            },
            {
              "id": 2871005,
              "postDate": "2024-06-14T02:23:33Z",
              "content": "<p>Lesson learned, i should have checked previous competition carefully. I saw several solutions that use postprocess that worked. I didn't use post process and previous data, a little bit pity.</p>",
              "rawMarkdown": "Lesson learned, i should have checked previous competition carefully. I saw several solutions that use postprocess that worked. I didn't use post process and previous data, a little bit pity.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2867896,
      "postDate": "2024-06-12T06:09:15.733Z",
      "content": "<p>Congrats for reaching Competition Grandmaster status 🎉 Huge achievement</p>",
      "rawMarkdown": "Congrats for reaching Competition Grandmaster status 🎉 Huge achievement",
      "votes": 1
    },
    {
      "id": 2866944,
      "postDate": "2024-06-11T15:25:39.187Z",
      "content": "<p>Congrats for your medal!</p>\n<p>Question about your signal approach, what is the intuition behind your reshaping? Why picking 2 channels instead of 3 for example?</p>",
      "rawMarkdown": "Congrats for your medal!\n\nQuestion about your signal approach, what is the intuition behind your reshaping? Why picking 2 channels instead of 3 for example?",
      "votes": 1,
      "replies": [
        {
          "id": 2867029,
          "postDate": "2024-06-11T15:50:15.747Z",
          "content": "<p>I want to make a 1 channel input ,but the input size is too big. <br>\nAnd that makes it too hard to do ensemble. <br>\nThen i reshape it as 2 channels( [123456]-&gt;[[135], [246]]). In some way, you can regard the input as 2  waves that with 16000hz sample rate.</p>",
          "rawMarkdown": "I want to make a 1 channel input ,but the input size is too big. \nAnd that makes it too hard to do ensemble. \nThen i reshape it as 2 channels( [123456]->[[135], [246]]). In some way, you can regard the input as 2  waves that with 16000hz sample rate.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2866892,
      "postDate": "2024-06-11T14:53:50.193Z",
      "content": "<p>Nice use of 1D-signals again! Also, congrats on becoming GM.</p>\n<p>Is there a reason why you use n_mels=512 and then resize to (256, 256) rather than just setting n_mels=256?</p>",
      "rawMarkdown": "Nice use of 1D-signals again! Also, congrats on becoming GM.\n\nIs there a reason why you use n_mels=512 and then resize to (256, 256) rather than just setting n_mels=256?",
      "votes": 1,
      "replies": [
        {
          "id": 2866923,
          "postDate": "2024-06-11T15:11:27.893Z",
          "content": "<p>For the spectrum model i use raw spectrum, but for the mix model, i use resized 256x256 for inference speed reason, otherwise, it would be timeout.</p>",
          "rawMarkdown": "For the spectrum model i use raw spectrum, but for the mix model, i use resized 256x256 for inference speed reason, otherwise, it would be timeout.",
          "votes": 1,
          "replies": [
            {
              "id": 2867030,
              "postDate": "2024-06-11T15:51:20.997Z",
              "content": "<p>Hmm okay, thanks for the reply.</p>\n<p>Why you did you not set <code>n_mels=256</code>? This should make the height match the resize shape?</p>\n<p>eg. </p>\n<pre><code>import torch\nimport torch.nn as nn\nimport torchaudio\n\nl= nn.Sequential(\n    torchaudio.transforms.MelSpectrogram(\n              32000,\n              =256, # &lt;- change this  256?\n              =0,\n              =16000,\n              =2048*2,\n              =512,\n              =,\n          ),\ntorchaudio.transforms.AmplitudeToDB(=80),\n)\nx= torch.rand(32_000)\n(l(x).shape)\n\n\n</code></pre>",
              "rawMarkdown": "Hmm okay, thanks for the reply.\n\nWhy you did you not set `n_mels=256`? This should make the height match the resize shape?\n\neg. \n\n```\nimport torch\nimport torch.nn as nn\nimport torchaudio\n\nl= nn.Sequential(\n    torchaudio.transforms.MelSpectrogram(\n              32000,\n              n_mels=256, # <- change this to 256?\n              f_min=0,\n              f_max=16000,\n              n_fft=2048*2,\n              hop_length=512,\n              normalized=True,\n          ),\ntorchaudio.transforms.AmplitudeToDB(top_db=80),\n)\nx= torch.rand(32_000*5)\nprint(l(x).shape)\n\n# resize here..\n```"
            },
            {
              "id": 2867652,
              "postDate": "2024-06-12T01:55:31.503Z",
              "content": "<p>Yeah， My models need raw data, spectrum512, spectrum256.<br>\nTherefor, the dataiter produce raw, spectrum512, spectrum256(resized by spectrum512). In this way i just need to load every ogg file one time to reduce the time cost. <br>\nI use n_mel=512 because i developed my spectrum model( input size 512,313) early and that is the best params for my model.</p>",
              "rawMarkdown": "Yeah， My models need raw data, spectrum512, spectrum256.\nTherefor, the dataiter produce raw, spectrum512, spectrum256(resized by spectrum512). In this way i just need to load every ogg file one time to reduce the time cost. \nI use n_mel=512 because i developed my spectrum model( input size 512,313) early and that is the best params for my model.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2866413,
      "postDate": "2024-06-11T10:14:01.903Z",
      "content": "<p>Great how you managed to train raw model that well, in my experiments it had bad CV so I didn't submit it</p>",
      "rawMarkdown": "Great how you managed to train raw model that well, in my experiments it had bad CV so I didn't submit it",
      "votes": 1
    },
    {
      "id": 2866141,
      "postDate": "2024-06-11T06:49:41.843Z",
      "content": "<p>Hi, I had a doubt like taking 5 random seconds from sample during training wont guarantee the bird noise in it right? If yes then how the model works?</p>",
      "rawMarkdown": "Hi, I had a doubt like taking 5 random seconds from sample during training wont guarantee the bird noise in it right? If yes then how the model works?",
      "votes": 1,
      "replies": [
        {
          "id": 2866174,
          "postDate": "2024-06-11T07:11:42.023Z",
          "content": "<p>I see some people train with first 5 sec data. But in my experiments, random sample 5 sec data, perform better.</p>",
          "rawMarkdown": "I see some people train with first 5 sec data. But in my experiments, random sample 5 sec data, perform better.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2866332,
      "postDate": "2024-06-11T09:15:34.257Z",
      "content": "<p>Solution seems so simple! How did you get such powerful single models using this approach? Were there any other tricks you used or parameters you optimized to get such high scoring models? Im so curious how such basic settings score so high. And congrats on GM rank!</p>",
      "rawMarkdown": "Solution seems so simple! How did you get such powerful single models using this approach? Were there any other tricks you used or parameters you optimized to get such high scoring models? Im so curious how such basic settings score so high. And congrats on GM rank!",
      "votes": 2,
      "replies": [
        {
          "id": 2866334,
          "postDate": "2024-06-11T09:16:23.080Z",
          "content": "<p>Would you be willing to share training code?</p>",
          "rawMarkdown": "Would you be willing to share training code?"
        },
        {
          "id": 2866398,
          "postDate": "2024-06-11T10:03:44.310Z",
          "content": "<p>Thanks. <br>\nThe model is quiet simple, but it's sensitive to random seed. I use lb to check model performance. Single model is not that powerful. But the ensemble of  raw signal and spectrum helps a lot.</p>",
          "rawMarkdown": "Thanks. \nThe model is quiet simple, but it's sensitive to random seed. I use lb to check model performance. Single model is not that powerful. But the ensemble of  raw signal and spectrum helps a lot.",
          "replies": [
            {
              "id": 2867017,
              "postDate": "2024-06-11T15:42:33.133Z",
              "content": "<p>What I’m asking is how did you get the spectrum only model to 0.70 public score when the baseline score given by most public notebook or discussion showed only 0.66 public score for spectrum only.</p>",
              "rawMarkdown": "What I’m asking is how did you get the spectrum only model to 0.70 public score when the baseline score given by most public notebook or discussion showed only 0.66 public score for spectrum only."
            }
          ]
        }
      ]
    },
    {
      "id": 2866143,
      "postDate": "2024-06-11T06:52:31.407Z",
      "content": "<p>Impressive!! Will you mind providing some more details about the raw signals model?<br>\nI think it is really a big breakthrough because current solutions like birdnet, bird-vocalization-classifier are all based on computer vision, thus solution with raw signal maybe cutting edge.</p>",
      "rawMarkdown": "Impressive!! Will you mind providing some more details about the raw signals model?\nI think it is really a big breakthrough because current solutions like birdnet, bird-vocalization-classifier are all based on computer vision, thus solution with raw signal maybe cutting edge.",
      "votes": 2,
      "replies": [
        {
          "id": 2866172,
          "postDate": "2024-06-11T07:10:29.340Z",
          "content": "<p>Like the code below,</p>\n<pre><code> = x.view(bs, -, )\n = torch.transpose(x, , )\n = x.view(bs, , -, )\n</code></pre>\n<p>then feeding to a 2d-CNN</p>\n<p>And i add some comment one the pse codes, i think you can get it now.</p>",
          "rawMarkdown": "Like the code below,\n```\nx = x.view(bs, -1, 2)\nx = torch.transpose(x, 2, 1)\nx = x.view(bs, 2, -1, 80)\n```\n\nthen feeding to a 2d-CNN\n\nAnd i add some comment one the pse codes, i think you can get it now.",
          "votes": 4
        },
        {
          "id": 2867154,
          "postDate": "2024-06-11T16:53:38.527Z",
          "content": "<p>I am not sure if it makes sense, but here is my thought since the model was trained on raw waves model might had learned to predict very well for the noisy data , it might have created some range pass that range it might be treating those features as an outlier, since the test data had too much noise it would had discarded all those unnecssary features upto some range based on training data, and since we normalize the value for mels we might not had those outliers , i am still confused tho :)  </p>",
          "rawMarkdown": "I am not sure if it makes sense, but here is my thought since the model was trained on raw waves model might had learned to predict very well for the noisy data , it might have created some range pass that range it might be treating those features as an outlier, since the test data had too much noise it would had discarded all those unnecssary features upto some range based on training data, and since we normalize the value for mels we might not had those outliers , i am still confused tho :)  "
        }
      ]
    },
    {
      "id": 2871578,
      "postDate": "2024-06-14T09:23:09.860Z",
      "content": "<p>Congratulations 👏</p>",
      "rawMarkdown": "Congratulations 👏"
    },
    {
      "id": 2869286,
      "postDate": "2024-06-13T02:09:33.057Z",
      "content": "<p>Congrats for becoming Competition Grandmaster!!!!!</p>",
      "rawMarkdown": "Congrats for becoming Competition Grandmaster!!!!!"
    },
    {
      "id": 2868710,
      "postDate": "2024-06-12T15:13:20.170Z",
      "content": "<p>Congrats for becoming a competition GM! Huge achievement!  Also impressive to see the relatively simple solution can achieve the great final result. </p>\n<p>Have the following questions: </p>\n<ol>\n<li><p>Why did you choose 5-sec clip during the training stage? In my experience, 10-sec clip from the middle of the train audio improves the cv score a bit, and also gives the model more samples to learn from my opinion. </p></li>\n<li><p>Why didn't you use more image/audio augmentations? </p></li>\n<li><p>Is there a reason for choosing n_fft=2048 * 2? That seems a lot more computation required than when n_fft=512, which some solutions adopt to use.</p></li>\n</ol>\n<p>Thanks again for sharing the solution!</p>",
      "rawMarkdown": "Congrats for becoming a competition GM! Huge achievement!  Also impressive to see the relatively simple solution can achieve the great final result. \n\nHave the following questions: \n\n1. Why did you choose 5-sec clip during the training stage? In my experience, 10-sec clip from the middle of the train audio improves the cv score a bit, and also gives the model more samples to learn from my opinion. \n\n2. Why didn't you use more image/audio augmentations? \n\n3. Is there a reason for choosing n_fft=2048 * 2? That seems a lot more computation required than when n_fft=512, which some solutions adopt to use.\n\nThanks again for sharing the solution!"
    },
    {
      "id": 2868227,
      "postDate": "2024-06-12T08:52:23.450Z",
      "content": "<p>Well done. The idea to mix raw and MEL spectrogram is great.<br>\nCould you give more details about the <code>openvino</code> part, thanks!</p>",
      "rawMarkdown": "Well done. The idea to mix raw and MEL spectrogram is great.\nCould you give more details about the `openvino` part, thanks!",
      "replies": [
        {
          "id": 2868243,
          "postDate": "2024-06-12T09:15:46.663Z",
          "content": "<p>Like most people, convert pytorch model to onnx, then use openvino do inference. But spectrum is produced by torchaudio. And I will share my notebook later.</p>",
          "rawMarkdown": "Like most people, convert pytorch model to onnx, then use openvino do inference. But spectrum is produced by torchaudio. And I will share my notebook later.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2868184,
      "postDate": "2024-06-12T08:31:25.813Z",
      "content": "<p>Congratulations! I've just started studying I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you.</p>",
      "rawMarkdown": "Congratulations! I've just started studying I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you."
    },
    {
      "id": 2867811,
      "postDate": "2024-06-12T05:15:59.517Z",
      "content": "<p>Congratulations on securing 5th Place in this competition. Also, thanks on sharing the solution details. </p>",
      "rawMarkdown": "Congratulations on securing 5th Place in this competition. Also, thanks on sharing the solution details. "
    },
    {
      "id": 2867231,
      "postDate": "2024-06-11T18:03:10.843Z",
      "content": "<p>Thanks for sharing and congratulations on the 5th place and GM title!</p>\n<p>I'm surprised that 1D models worked. We tried them early and discarded the idea because the performance was considerably worse than melspecs. Maybe the size of the test made a difference (only part of the data for 1 fold and 10 epochs). A reminder to revisit early decisions.</p>\n<p>Could you please elaborate a bit on the approach you followed to evaluate models and define the weights for the ensemble? This was a struggle for us and I wonder how other teams approached it.</p>",
      "rawMarkdown": "Thanks for sharing and congratulations on the 5th place and GM title!\n\nI'm surprised that 1D models worked. We tried them early and discarded the idea because the performance was considerably worse than melspecs. Maybe the size of the test made a difference (only part of the data for 1 fold and 10 epochs). A reminder to revisit early decisions.\n\nCould you please elaborate a bit on the approach you followed to evaluate models and define the weights for the ensemble? This was a struggle for us and I wonder how other teams approached it.",
      "replies": [
        {
          "id": 2868149,
          "postDate": "2024-06-12T08:00:23.920Z",
          "content": "<p>I didn't find a good way to do local cv, and raw signal model cv is lower than spectrum model. <br>\nI use 0.5,0.4,0.1 as ensemble weights. Because raw signal model score is about the same with spectrum model, and the mix model is a combine of them, use a small weight here. </p>",
          "rawMarkdown": "I didn't find a good way to do local cv, and raw signal model cv is lower than spectrum model. \nI use 0.5,0.4,0.1 as ensemble weights. Because raw signal model score is about the same with spectrum model, and the mix model is a combine of them, use a small weight here. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2866485,
      "postDate": "2024-06-11T11:09:11.620Z",
      "content": "<p>Congratulations ! Would you be publishing your inference pipeline? I'd be very interested to see how you optimised MelSpec layers with openvino, I failed at this. </p>",
      "rawMarkdown": "Congratulations ! Would you be publishing your inference pipeline? I'd be very interested to see how you optimised MelSpec layers with openvino, I failed at this. ",
      "replies": [
        {
          "id": 2866495,
          "postDate": "2024-06-11T11:11:46.423Z",
          "content": "<p>I will share it later. However I met the same problem, I use torchaudio to produce the spectrum in dataiter, the layer not included in the model.</p>",
          "rawMarkdown": "I will share it later. However I met the same problem, I use torchaudio to produce the spectrum in dataiter, the layer not included in the model."
        },
        {
          "id": 2866574,
          "postDate": "2024-06-11T11:46:09.813Z",
          "content": "<p>Try my last year <a href=\"https://www.kaggle.com/code/aphysict/fork-of-birdclef-submit-rewrite-onnx\" target=\"_blank\">notebook</a></p>",
          "rawMarkdown": "Try my last year [notebook](https://www.kaggle.com/code/aphysict/fork-of-birdclef-submit-rewrite-onnx)",
          "votes": 1,
          "replies": [
            {
              "id": 2868130,
              "postDate": "2024-06-12T07:50:29.093Z",
              "content": "<p>Thanks, I will go through your NB. </p>",
              "rawMarkdown": "Thanks, I will go through your NB. "
            }
          ]
        }
      ]
    },
    {
      "id": 3146440,
      "postDate": "2025-03-10T21:39:57.853Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2875213,
      "postDate": "2024-06-17T00:28:55.100Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2866694,
      "postDate": "2024-06-11T13:02:35.197Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2866062,
      "author_name": "Donghui Zhang",
      "author_url": "",
      "post_date": "2024-06-11T06:04:18.460000",
      "content": "<p>Regarding the \"raw model\", how did you deal with it to achieve good results? I rarely see people using raw signals in Bird competitions.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2866242,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-11T08:08:36.480000",
          "content": "<p>I add some comments one the pse codes, i think you can get it now.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2869294,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-06-13T02:25:50.187000",
      "content": "<p>Good work. I was betting you'd use packed raw waveform as you did in HMS competition!</p>\n<p>Congrats on becoming a GM. It was a matter of time given what you did in HMS.</p>\n<p>Did you compare openvino accuracy with pytorch models?  Asking because we found that it loses some accuracy and we didn't use it as a result.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2869391,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-13T04:32:40.203000",
          "content": "<p>Thanks.</p>\n<p>I noticed that you'd mentioned the precision loss with openvino. <br>\nI manully check some of the outputs of pytorch, onnxruntime, and openvino, and notice very small differences, but didn't try to sub with onnxruntime. By the way, as the the score is rank relavent, i use logits as final output. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2870282,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-06-13T13:48:26.987000",
              "content": "<p>I used logits too until we found that smoothing worked better when applied to probabilities.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2871005,
              "author_name": "coolz",
              "author_url": "",
              "post_date": "2024-06-14T02:23:33",
              "content": "<p>Lesson learned, i should have checked previous competition carefully. I saw several solutions that use postprocess that worked. I didn't use post process and previous data, a little bit pity.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2867896,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2024-06-12T06:09:15.733000",
      "content": "<p>Congrats for reaching Competition Grandmaster status 🎉 Huge achievement</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2866944,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2024-06-11T15:25:39.187000",
      "content": "<p>Congrats for your medal!</p>\n<p>Question about your signal approach, what is the intuition behind your reshaping? Why picking 2 channels instead of 3 for example?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2867029,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-11T15:50:15.747000",
          "content": "<p>I want to make a 1 channel input ,but the input size is too big. <br>\nAnd that makes it too hard to do ensemble. <br>\nThen i reshape it as 2 channels( [123456]-&gt;[[135], [246]]). In some way, you can regard the input as 2  waves that with 16000hz sample rate.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2866892,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-06-11T14:53:50.193000",
      "content": "<p>Nice use of 1D-signals again! Also, congrats on becoming GM.</p>\n<p>Is there a reason why you use n_mels=512 and then resize to (256, 256) rather than just setting n_mels=256?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2866923,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-11T15:11:27.893000",
          "content": "<p>For the spectrum model i use raw spectrum, but for the mix model, i use resized 256x256 for inference speed reason, otherwise, it would be timeout.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2867030,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-06-11T15:51:20.997000",
              "content": "<p>Hmm okay, thanks for the reply.</p>\n<p>Why you did you not set <code>n_mels=256</code>? This should make the height match the resize shape?</p>\n<p>eg. </p>\n<pre><code>import torch\nimport torch.nn as nn\nimport torchaudio\n\nl= nn.Sequential(\n    torchaudio.transforms.MelSpectrogram(\n              32000,\n              =256, # &lt;- change this  256?\n              =0,\n              =16000,\n              =2048*2,\n              =512,\n              =,\n          ),\ntorchaudio.transforms.AmplitudeToDB(=80),\n)\nx= torch.rand(32_000)\n(l(x).shape)\n\n\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2867652,
              "author_name": "coolz",
              "author_url": "",
              "post_date": "2024-06-12T01:55:31.503000",
              "content": "<p>Yeah， My models need raw data, spectrum512, spectrum256.<br>\nTherefor, the dataiter produce raw, spectrum512, spectrum256(resized by spectrum512). In this way i just need to load every ogg file one time to reduce the time cost. <br>\nI use n_mel=512 because i developed my spectrum model( input size 512,313) early and that is the best params for my model.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2866413,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-06-11T10:14:01.903000",
      "content": "<p>Great how you managed to train raw model that well, in my experiments it had bad CV so I didn't submit it</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2866141,
      "author_name": "Ayush Solanki",
      "author_url": "",
      "post_date": "2024-06-11T06:49:41.843000",
      "content": "<p>Hi, I had a doubt like taking 5 random seconds from sample during training wont guarantee the bird noise in it right? If yes then how the model works?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2866174,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-11T07:11:42.023000",
          "content": "<p>I see some people train with first 5 sec data. But in my experiments, random sample 5 sec data, perform better.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2866332,
      "author_name": "snehal",
      "author_url": "",
      "post_date": "2024-06-11T09:15:34.257000",
      "content": "<p>Solution seems so simple! How did you get such powerful single models using this approach? Were there any other tricks you used or parameters you optimized to get such high scoring models? Im so curious how such basic settings score so high. And congrats on GM rank!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2866334,
          "author_name": "snehal",
          "author_url": "",
          "post_date": "2024-06-11T09:16:23.080000",
          "content": "<p>Would you be willing to share training code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2866398,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-11T10:03:44.310000",
          "content": "<p>Thanks. <br>\nThe model is quiet simple, but it's sensitive to random seed. I use lb to check model performance. Single model is not that powerful. But the ensemble of  raw signal and spectrum helps a lot.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2867017,
              "author_name": "snehal",
              "author_url": "",
              "post_date": "2024-06-11T15:42:33.133000",
              "content": "<p>What I’m asking is how did you get the spectrum only model to 0.70 public score when the baseline score given by most public notebook or discussion showed only 0.66 public score for spectrum only.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2866143,
      "author_name": "RihanPiggy",
      "author_url": "",
      "post_date": "2024-06-11T06:52:31.407000",
      "content": "<p>Impressive!! Will you mind providing some more details about the raw signals model?<br>\nI think it is really a big breakthrough because current solutions like birdnet, bird-vocalization-classifier are all based on computer vision, thus solution with raw signal maybe cutting edge.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2866172,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-11T07:10:29.340000",
          "content": "<p>Like the code below,</p>\n<pre><code> = x.view(bs, -, )\n = torch.transpose(x, , )\n = x.view(bs, , -, )\n</code></pre>\n<p>then feeding to a 2d-CNN</p>\n<p>And i add some comment one the pse codes, i think you can get it now.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2867154,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2024-06-11T16:53:38.527000",
          "content": "<p>I am not sure if it makes sense, but here is my thought since the model was trained on raw waves model might had learned to predict very well for the noisy data , it might have created some range pass that range it might be treating those features as an outlier, since the test data had too much noise it would had discarded all those unnecssary features upto some range based on training data, and since we normalize the value for mels we might not had those outliers , i am still confused tho :)  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2871578,
      "author_name": "Muhammad Azeem",
      "author_url": "",
      "post_date": "2024-06-14T09:23:09.860000",
      "content": "<p>Congratulations 👏</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2869286,
      "author_name": "Aaditya Porwal",
      "author_url": "",
      "post_date": "2024-06-13T02:09:33.057000",
      "content": "<p>Congrats for becoming Competition Grandmaster!!!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2868710,
      "author_name": "faithk7u",
      "author_url": "",
      "post_date": "2024-06-12T15:13:20.170000",
      "content": "<p>Congrats for becoming a competition GM! Huge achievement!  Also impressive to see the relatively simple solution can achieve the great final result. </p>\n<p>Have the following questions: </p>\n<ol>\n<li><p>Why did you choose 5-sec clip during the training stage? In my experience, 10-sec clip from the middle of the train audio improves the cv score a bit, and also gives the model more samples to learn from my opinion. </p></li>\n<li><p>Why didn't you use more image/audio augmentations? </p></li>\n<li><p>Is there a reason for choosing n_fft=2048 * 2? That seems a lot more computation required than when n_fft=512, which some solutions adopt to use.</p></li>\n</ol>\n<p>Thanks again for sharing the solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2868227,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2024-06-12T08:52:23.450000",
      "content": "<p>Well done. The idea to mix raw and MEL spectrogram is great.<br>\nCould you give more details about the <code>openvino</code> part, thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2868243,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-12T09:15:46.663000",
          "content": "<p>Like most people, convert pytorch model to onnx, then use openvino do inference. But spectrum is produced by torchaudio. And I will share my notebook later.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2868184,
      "author_name": "YoungJun_Player",
      "author_url": "",
      "post_date": "2024-06-12T08:31:25.813000",
      "content": "<p>Congratulations! I've just started studying I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2867811,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-06-12T05:15:59.517000",
      "content": "<p>Congratulations on securing 5th Place in this competition. Also, thanks on sharing the solution details. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2867231,
      "author_name": "vialactea",
      "author_url": "",
      "post_date": "2024-06-11T18:03:10.843000",
      "content": "<p>Thanks for sharing and congratulations on the 5th place and GM title!</p>\n<p>I'm surprised that 1D models worked. We tried them early and discarded the idea because the performance was considerably worse than melspecs. Maybe the size of the test made a difference (only part of the data for 1 fold and 10 epochs). A reminder to revisit early decisions.</p>\n<p>Could you please elaborate a bit on the approach you followed to evaluate models and define the weights for the ensemble? This was a struggle for us and I wonder how other teams approached it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2868149,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-12T08:00:23.920000",
          "content": "<p>I didn't find a good way to do local cv, and raw signal model cv is lower than spectrum model. <br>\nI use 0.5,0.4,0.1 as ensemble weights. Because raw signal model score is about the same with spectrum model, and the mix model is a combine of them, use a small weight here. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2866485,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2024-06-11T11:09:11.620000",
      "content": "<p>Congratulations ! Would you be publishing your inference pipeline? I'd be very interested to see how you optimised MelSpec layers with openvino, I failed at this. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2866495,
          "author_name": "coolz",
          "author_url": "",
          "post_date": "2024-06-11T11:11:46.423000",
          "content": "<p>I will share it later. However I met the same problem, I use torchaudio to produce the spectrum in dataiter, the layer not included in the model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2866574,
          "author_name": "Aphysict",
          "author_url": "",
          "post_date": "2024-06-11T11:46:09.813000",
          "content": "<p>Try my last year <a href=\"https://www.kaggle.com/code/aphysict/fork-of-birdclef-submit-rewrite-onnx\" target=\"_blank\">notebook</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2868130,
              "author_name": "Phaedrus",
              "author_url": "",
              "post_date": "2024-06-12T07:50:29.093000",
              "content": "<p>Thanks, I will go through your NB. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3146440,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-10T21:39:57.853000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2875213,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-17T00:28:55.100000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2866694,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-11T13:02:35.197000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2866034": "\n\nFirst of all, thanks to the organizers for hosting this competition.\nAnd congrats to all winners.\n\n\n### Data\n2024 data\n\n### Model\n| model                          | backbone       | weights     | public   | private  |\n|--------------------------------|----------------|-------------|----------|----------|\n| raw signale                    | hgnetb0        | 3/5 folds   | 0.720354 | 0.666671 |\n| spectrum                       | efficientnetb0 | 1/5 folds   | 0.708198    | 0.672360    |\n| raw+spectrum                   | hgnetb0+effb0  | 1/5 folds   | 0.698408 | 0.660503 |\n| ensemble                       | ...            | 0.5+0.4+0.1 | 0.743960 | 0.687173 |\n\n+ 1. raw signal model\n```\n# x bsx160000\n\n#bsx80000x2\nx = x.view(bs, -1, 2)           \n#bsx2x80000\nx = torch.transpose(x, 2, 1)    \n#bsx2x1000x80\nx = x.view(bs, 2, -1, 80)         \nfeature = self.backbone(x)   \n```\n+ 2. spectrum model\n```\ntorchaudio.transforms.MelSpectrogram(\n                32000,\n                n_mels=512,\n                f_min=0,\n                f_max=16000,\n                n_fft=2048*2,\n                hop_length=512,\n                normalized=True,\n            ),\n\ntorchaudio.transforms.AmplitudeToDB(top_db=80)\n```\n+ 3. mix model\n\n```commandline\nraw_f  = self.raw_model(x)\nspec_f = self.spec_model(x)\nfeature = torch.cat([raw_f,spec_f],dim=1)\nx=self.fc(feature)\n\n## spectrum is resized to 256x256 for fast speed, \n## other params for mel spectrum is the same with spectrm model.\n```\n### Preprocess\n+ 1. wave/max(wave) if max(wave)>1\n### Augmentation\n+ XY cut out for spectrum model\n### Train\n+ 1. Random sample 5 seconds data for trainning, \nfirst 5 seconds data for validation.\n+ 2. Mixup p=1.\n+ 3. BCEWithLogitsLoss\n\n### inference\n+ openvino\n\n\n### extra\nI make a post to discussion the method feeding raw signal to vision model. Here is the link [good luck](https://www.kaggle.com/competitions/birdclef-2024/discussion/511763)\n\n\ninference [notebook](https://www.kaggle.com/code/cooolz/5nd-solution)\n",
    "2866062": "Regarding the \"raw model\", how did you deal with it to achieve good results? I rarely see people using raw signals in Bird competitions.",
    "2869294": "Good work. I was betting you'd use packed raw waveform as you did in HMS competition!\n\nCongrats on becoming a GM. It was a matter of time given what you did in HMS.\n\nDid you compare openvino accuracy with pytorch models?  Asking because we found that it loses some accuracy and we didn't use it as a result.",
    "2867896": "Congrats for reaching Competition Grandmaster status 🎉 Huge achievement",
    "2866944": "Congrats for your medal!\n\nQuestion about your signal approach, what is the intuition behind your reshaping? Why picking 2 channels instead of 3 for example?",
    "2866892": "Nice use of 1D-signals again! Also, congrats on becoming GM.\n\nIs there a reason why you use n_mels=512 and then resize to (256, 256) rather than just setting n_mels=256?",
    "2866413": "Great how you managed to train raw model that well, in my experiments it had bad CV so I didn't submit it",
    "2866141": "Hi, I had a doubt like taking 5 random seconds from sample during training wont guarantee the bird noise in it right? If yes then how the model works?",
    "2866332": "Solution seems so simple! How did you get such powerful single models using this approach? Were there any other tricks you used or parameters you optimized to get such high scoring models? Im so curious how such basic settings score so high. And congrats on GM rank!",
    "2866143": "Impressive!! Will you mind providing some more details about the raw signals model?\nI think it is really a big breakthrough because current solutions like birdnet, bird-vocalization-classifier are all based on computer vision, thus solution with raw signal maybe cutting edge.",
    "2871578": "Congratulations 👏",
    "2869286": "Congrats for becoming Competition Grandmaster!!!!!",
    "2868710": "Congrats for becoming a competition GM! Huge achievement!  Also impressive to see the relatively simple solution can achieve the great final result. \n\nHave the following questions: \n\n1. Why did you choose 5-sec clip during the training stage? In my experience, 10-sec clip from the middle of the train audio improves the cv score a bit, and also gives the model more samples to learn from my opinion. \n\n2. Why didn't you use more image/audio augmentations? \n\n3. Is there a reason for choosing n_fft=2048 * 2? That seems a lot more computation required than when n_fft=512, which some solutions adopt to use.\n\nThanks again for sharing the solution!",
    "2868227": "Well done. The idea to mix raw and MEL spectrogram is great.\nCould you give more details about the `openvino` part, thanks!",
    "2868184": "Congratulations! I've just started studying I feel that I can grow quickly with the help of people like you, who are willing to discuss, share, and provide guidance. Thank you.",
    "2867811": "Congratulations on securing 5th Place in this competition. Also, thanks on sharing the solution details. ",
    "2867231": "Thanks for sharing and congratulations on the 5th place and GM title!\n\nI'm surprised that 1D models worked. We tried them early and discarded the idea because the performance was considerably worse than melspecs. Maybe the size of the test made a difference (only part of the data for 1 fold and 10 epochs). A reminder to revisit early decisions.\n\nCould you please elaborate a bit on the approach you followed to evaluate models and define the weights for the ensemble? This was a struggle for us and I wonder how other teams approached it.",
    "2866485": "Congratulations ! Would you be publishing your inference pipeline? I'd be very interested to see how you optimised MelSpec layers with openvino, I failed at this. ",
    "3146440": "",
    "2875213": "",
    "2866694": ""
  }
}