{
  "id": 327044,
  "title": "5th place solution",
  "url": "/competitions/birdclef-2022/writeups/common-kestrel-5th-place-solution",
  "author_name": "",
  "post_date": "2022-05-25T13:37:26.687Z",
  "votes": 44,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Many thanks to Kaggle and Cornell Lab of Ornithology for hosting such an interesting competition.</p>\n<p>My solution is a reimplementation of the BirdCLEF 2021 <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution</a>.<br>\nSpecial thanks to the new baseline team for publishing the amazing solution. And thanks to <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> for sharing <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">the great baseline!</a></p>\n<h1>Model</h1>\n<p>My final submission is an ensemble of 9 models of different seed/fold/backbone, each model using last year's 2nd place solution method.</p>\n<p>The backbones were:</p>\n<ul>\n<li>4x <code>eca_nfnet_l0</code></li>\n<li>2x <code>tf_efficientnetv2_s_in21k</code></li>\n<li>2x <code>resnet34</code></li>\n<li>1x <code>convnext_tiny</code></li>\n</ul>\n<h1>Oversampling</h1>\n<p>To increase the number of files for minority classes(N&lt;20), I split the training files by hand.<br>\nFor example, <code>maupar</code> appears only in one file, but the file contains many different types of its songs and calls.<br>\nI cut it into segments of 10-30 seconds using a waveform editor (<a href=\"https://www.audacityteam.org/\" target=\"_blank\">Audacity</a>) and created multiple training data from a single file.</p>\n<p>The cut audio files were further augmented by applying effects such as noise reduction/reverb/gain/etc. to each segment in the waveform editor. Ultimately, I created 5-20 additional sample files per target minority class.</p>\n<h1>Training</h1>\n<ul>\n<li>Melspec<ol>\n<li>window_size=1024, hop_size=320, fmin=50, fmax=14000, power=2, mel_bins=64, top_db=80 (nfnet, effnet, convnext)</li>\n<li>window_size=2048, hop_size=512, fmin=16, fmax=16386, power=2, mel_bins=256, top_db=80 (resnet)</li></ol></li>\n<li>Data Augmentation(Waveform)<ul>\n<li>Background Noise (2020 nocall, 2021 nocall, freefield1010)</li>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>NoiseInjection</li>\n<li>RandomVolume</li>\n<li>TimeShift</li></ul></li>\n<li>Data Augmentation(Image)<ul>\n<li>SpecAug</li>\n<li>CutOut</li>\n<li>Lowpass</li>\n<li>TranslateY (shift in freq dim, simulated pitch shift)</li></ul></li>\n<li>Optimizer<ul>\n<li>AdamW, LR=1e-3, weight decay=1e-5, Cosine Anearling with warmup</li></ul></li>\n<li>Loss<ul>\n<li>BCEWithLogitsLoss</li>\n<li><a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66?scriptVersionId=88985374\" target=\"_blank\">Weighted loss by rating</a></li></ul></li>\n<li>25-30 epochs</li>\n<li>Training target is the union of primary and secondary labels.</li>\n<li>All backbones were <strong>pretrained on BirdCLEF2021</strong> data.</li>\n<li><strong>micro-f1@0.1</strong> for <code>scored_birds</code> was used for validation. Although a bit odd, the correlation between this score and Public/Private LB was not bad.</li>\n</ul>\n<h1>Post Processing</h1>\n<ul>\n<li>Probability averaged over previous and next chunks.</li>\n<li>Correct thresholds using <a href=\"https://github.com/ChristofHenkel/kaggle-birdclef2021-2nd-place/blob/main/configs/pp_binary_ext3_1.py\" target=\"_blank\">call/nocall binary classifier</a> probabilities.</li>\n<li>Threshold optimization by percentile per clesses.</li>\n</ul>\n<pre><code># thresholds per class\nthreshold = pd.Series(np.percentile(test_df[SCORED_BIRDS].values, 90, axis=0), index=SCORED_BIRDS)\n</code></pre>\n<h1>What Did Not Work</h1>\n<ul>\n<li>SED models - could not make it past 0.72 for me</li>\n<li>CoordConv</li>\n<li>MC Dropout</li>\n<li>Label Smoothing</li>\n<li>Undersampling for majority classes</li>\n<li>ViT / Swin backbone</li>\n</ul>",
  "messages": [
    {
      "id": "1800967",
      "postDate": "05/25/2022 10:32:09",
      "content": "<p>Many thanks to Kaggle and Cornell Lab of Ornithology for hosting such an interesting competition.</p>\n<p>My solution is a reimplementation of the BirdCLEF 2021 <a href=\"https://www.kaggle.com/competitions/birdclef-2021/discussion/243463\" target=\"_blank\">2nd place solution</a>.<br>\nSpecial thanks to the new baseline team for publishing the amazing solution. And thanks to <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> for sharing <a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66\" target=\"_blank\">the great baseline!</a></p>\n<h1>Model</h1>\n<p>My final submission is an ensemble of 9 models of different seed/fold/backbone, each model using last year's 2nd place solution method.</p>\n<p>The backbones were:</p>\n<ul>\n<li>4x <code>eca_nfnet_l0</code></li>\n<li>2x <code>tf_efficientnetv2_s_in21k</code></li>\n<li>2x <code>resnet34</code></li>\n<li>1x <code>convnext_tiny</code></li>\n</ul>\n<h1>Oversampling</h1>\n<p>To increase the number of files for minority classes(N&lt;20), I split the training files by hand.<br>\nFor example, <code>maupar</code> appears only in one file, but the file contains many different types of its songs and calls.<br>\nI cut it into segments of 10-30 seconds using a waveform editor (<a href=\"https://www.audacityteam.org/\" target=\"_blank\">Audacity</a>) and created multiple training data from a single file.</p>\n<p>The cut audio files were further augmented by applying effects such as noise reduction/reverb/gain/etc. to each segment in the waveform editor. Ultimately, I created 5-20 additional sample files per target minority class.</p>\n<h1>Training</h1>\n<ul>\n<li>Melspec<ol>\n<li>window_size=1024, hop_size=320, fmin=50, fmax=14000, power=2, mel_bins=64, top_db=80 (nfnet, effnet, convnext)</li>\n<li>window_size=2048, hop_size=512, fmin=16, fmax=16386, power=2, mel_bins=256, top_db=80 (resnet)</li></ol></li>\n<li>Data Augmentation(Waveform)<ul>\n<li>Background Noise (2020 nocall, 2021 nocall, freefield1010)</li>\n<li>GaussianNoise</li>\n<li>PinkNoise</li>\n<li>NoiseInjection</li>\n<li>RandomVolume</li>\n<li>TimeShift</li></ul></li>\n<li>Data Augmentation(Image)<ul>\n<li>SpecAug</li>\n<li>CutOut</li>\n<li>Lowpass</li>\n<li>TranslateY (shift in freq dim, simulated pitch shift)</li></ul></li>\n<li>Optimizer<ul>\n<li>AdamW, LR=1e-3, weight decay=1e-5, Cosine Anearling with warmup</li></ul></li>\n<li>Loss<ul>\n<li>BCEWithLogitsLoss</li>\n<li><a href=\"https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66?scriptVersionId=88985374\" target=\"_blank\">Weighted loss by rating</a></li></ul></li>\n<li>25-30 epochs</li>\n<li>Training target is the union of primary and secondary labels.</li>\n<li>All backbones were <strong>pretrained on BirdCLEF2021</strong> data.</li>\n<li><strong>micro-f1@0.1</strong> for <code>scored_birds</code> was used for validation. Although a bit odd, the correlation between this score and Public/Private LB was not bad.</li>\n</ul>\n<h1>Post Processing</h1>\n<ul>\n<li>Probability averaged over previous and next chunks.</li>\n<li>Correct thresholds using <a href=\"https://github.com/ChristofHenkel/kaggle-birdclef2021-2nd-place/blob/main/configs/pp_binary_ext3_1.py\" target=\"_blank\">call/nocall binary classifier</a> probabilities.</li>\n<li>Threshold optimization by percentile per clesses.</li>\n</ul>\n<pre><code># thresholds per class\nthreshold = pd.Series(np.percentile(test_df[SCORED_BIRDS].values, 90, axis=0), index=SCORED_BIRDS)\n</code></pre>\n<h1>What Did Not Work</h1>\n<ul>\n<li>SED models - could not make it past 0.72 for me</li>\n<li>CoordConv</li>\n<li>MC Dropout</li>\n<li>Label Smoothing</li>\n<li>Undersampling for majority classes</li>\n<li>ViT / Swin backbone</li>\n</ul>",
      "rawMarkdown": "Many thanks to Kaggle and Cornell Lab of Ornithology for hosting such an interesting competition.\n\nMy solution is a reimplementation of the BirdCLEF 2021 [2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463).\nSpecial thanks to the new baseline team for publishing the amazing solution. And thanks to @julian3833 for sharing [the great baseline!](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66)\n\n# Model\nMy final submission is an ensemble of 9 models of different seed/fold/backbone, each model using last year's 2nd place solution method.\n\nThe backbones were:\n* 4x `eca_nfnet_l0`\n* 2x `tf_efficientnetv2_s_in21k`\n* 2x `resnet34`\n* 1x `convnext_tiny`\n\n# Oversampling\n\nTo increase the number of files for minority classes(N<20), I split the training files by hand.\nFor example, `maupar` appears only in one file, but the file contains many different types of its songs and calls.\nI cut it into segments of 10-30 seconds using a waveform editor ([Audacity](https://www.audacityteam.org/)) and created multiple training data from a single file.\n\nThe cut audio files were further augmented by applying effects such as noise reduction/reverb/gain/etc. to each segment in the waveform editor. Ultimately, I created 5-20 additional sample files per target minority class.\n\n\n# Training\n\n* Melspec\n\t1. window_size=1024, hop_size=320, fmin=50, fmax=14000, power=2, mel_bins=64, top_db=80 (nfnet, effnet, convnext)\n\t2. window_size=2048, hop_size=512, fmin=16, fmax=16386, power=2, mel_bins=256, top_db=80 (resnet)\n* Data Augmentation(Waveform)\n\t* Background Noise (2020 nocall, 2021 nocall, freefield1010)\n\t* GaussianNoise\n\t* PinkNoise\n\t* NoiseInjection\n\t* RandomVolume\n\t* TimeShift\n* Data Augmentation(Image)\n\t* SpecAug\n\t* CutOut\n\t* Lowpass\n\t* TranslateY (shift in freq dim, simulated pitch shift)\n* Optimizer\n\t* AdamW, LR=1e-3, weight decay=1e-5, Cosine Anearling with warmup\n* Loss\n\t* BCEWithLogitsLoss\n\t* [Weighted loss by rating](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66?scriptVersionId=88985374)\n* 25-30 epochs\n* Training target is the union of primary and secondary labels.\n* All backbones were **pretrained on BirdCLEF2021** data.\n* **micro-f1@0.1** for `scored_birds` was used for validation. Although a bit odd, the correlation between this score and Public/Private LB was not bad.\n\n\n# Post Processing\n\n* Probability averaged over previous and next chunks.\n* Correct thresholds using [call/nocall binary classifier](https://github.com/ChristofHenkel/kaggle-birdclef2021-2nd-place/blob/main/configs/pp_binary_ext3_1.py) probabilities.\n* Threshold optimization by percentile per clesses.\n\n```python\n# thresholds per class\nthreshold = pd.Series(np.percentile(test_df[SCORED_BIRDS].values, 90, axis=0), index=SCORED_BIRDS)\n```\n\n# What Did Not Work\n\n* SED models - could not make it past 0.72 for me\n* CoordConv\n* MC Dropout\n* Label Smoothing\n* Undersampling for majority classes\n* ViT / Swin backbone",
      "votes": null
    },
    {
      "id": "1801009",
      "postDate": "05/25/2022 11:01:08",
      "content": "<p>Congrats on your solo gold and thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats on your solo gold and thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "1801653",
      "postDate": "05/26/2022 02:19:38",
      "content": "<p>Congrats on the solo gold!!</p>",
      "rawMarkdown": "Congrats on the solo gold!!",
      "votes": null
    },
    {
      "id": "1801669",
      "postDate": "05/26/2022 03:00:12",
      "content": "<p>Congrats. And you became competition master!</p>",
      "rawMarkdown": "Congrats. And you became competition master!",
      "votes": null
    },
    {
      "id": "1801708",
      "postDate": "05/26/2022 04:26:07",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>, and congratulations on becoming a Competitions Master too!</p>",
      "rawMarkdown": "Thanks @shinmurashinmura, and congratulations on becoming a Competitions Master too!",
      "votes": null
    },
    {
      "id": "1801869",
      "postDate": "05/26/2022 08:17:38",
      "content": "<p>Thank you. We are competition master.</p>",
      "rawMarkdown": "Thank you. We are competition master.",
      "votes": null
    },
    {
      "id": "1802192",
      "postDate": "05/26/2022 14:40:23",
      "content": "<p>Congrats! Like with 3rd place I am glad that our last year's solution was still working so well here. Thanks for the writeup.</p>",
      "rawMarkdown": "Congrats! Like with 3rd place I am glad that our last year's solution was still working so well here. Thanks for the writeup.",
      "votes": null
    },
    {
      "id": "1802260",
      "postDate": "05/26/2022 15:46:17",
      "content": "<p>Congratulations! 🎉🎉</p>",
      "rawMarkdown": "Congratulations! 🎉🎉",
      "votes": null
    },
    {
      "id": "1802340",
      "postDate": "05/26/2022 16:31:12",
      "content": "<p>Congrats for this result!</p>\n<p>I also tried the convext family but could not train it. I tried resizing the images to 224 x 224 (if I remember it well) in the albumentations step but it still did not solve the size mismatch error in layers. Did you make any pre-process or resizing step to be able to use convnext? I have never used this architecture before. I tried to use it as encoder to subsequent SED model, it did not work…</p>",
      "rawMarkdown": "Congrats for this result!\n\nI also tried the convext family but could not train it. I tried resizing the images to 224 x 224 (if I remember it well) in the albumentations step but it still did not solve the size mismatch error in layers. Did you make any pre-process or resizing step to be able to use convnext? I have never used this architecture before. I tried to use it as encoder to subsequent SED model, it did not work...",
      "votes": null
    },
    {
      "id": "1802714",
      "postDate": "05/27/2022 04:33:46",
      "content": "<p>Thanks!</p>\n<p>This is my implementation, but I don't think I did anything special other than following <a href=\"https://rwightman.github.io/pytorch-image-models/feature_extraction/\" target=\"_blank\">timm's style</a>.</p>\n<pre><code>import timm\nimport torch\n\n# create convnet\nbackbone = timm.create_model(\n                \"convnext_tiny\",\n                pretrained=True,\n                num_classes=0,\n                global_pool=\"\",\n                in_chans=1,\n                drop_path_rate=0.1,\n            )\nnum_features = backbone.num_features\n\n# inference\nwith torch.inference_mode():\n    mel = torch.rand((1,1,64,500))  # B C H W\n    y_pred = backbone(mel)  # =&gt; (1, 768, 2, 15)\n</code></pre>",
      "rawMarkdown": "Thanks!\n\nThis is my implementation, but I don't think I did anything special other than following [timm's style](https://rwightman.github.io/pytorch-image-models/feature_extraction/).\n\n```python\nimport timm\nimport torch\n\n# create convnet\nbackbone = timm.create_model(\n                \"convnext_tiny\",\n                pretrained=True,\n                num_classes=0,\n                global_pool=\"\",\n                in_chans=1,\n                drop_path_rate=0.1,\n            )\nnum_features = backbone.num_features\n\n# inference\nwith torch.inference_mode():\n    mel = torch.rand((1,1,64,500))  # B C H W\n    y_pred = backbone(mel)  # => (1, 768, 2, 15)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801009,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "05/25/2022 11:01:08",
      "content": "<p>Congrats on your solo gold and thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801653,
      "author_name": "shotaosaki",
      "author_url": "",
      "post_date": "05/26/2022 02:19:38",
      "content": "<p>Congrats on the solo gold!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1801669,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "05/26/2022 03:00:12",
      "content": "<p>Congrats. And you became competition master!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801708,
          "author_name": "yokuyama",
          "author_url": "",
          "post_date": "05/26/2022 04:26:07",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a>, and congratulations on becoming a Competitions Master too!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801869,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "05/26/2022 08:17:38",
          "content": "<p>Thank you. We are competition master.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1802192,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "05/26/2022 14:40:23",
      "content": "<p>Congrats! Like with 3rd place I am glad that our last year's solution was still working so well here. Thanks for the writeup.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1802260,
      "author_name": "julian3833",
      "author_url": "",
      "post_date": "05/26/2022 15:46:17",
      "content": "<p>Congratulations! 🎉🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1802340,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "05/26/2022 16:31:12",
      "content": "<p>Congrats for this result!</p>\n<p>I also tried the convext family but could not train it. I tried resizing the images to 224 x 224 (if I remember it well) in the albumentations step but it still did not solve the size mismatch error in layers. Did you make any pre-process or resizing step to be able to use convnext? I have never used this architecture before. I tried to use it as encoder to subsequent SED model, it did not work…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1802714,
          "author_name": "yokuyama",
          "author_url": "",
          "post_date": "05/27/2022 04:33:46",
          "content": "<p>Thanks!</p>\n<p>This is my implementation, but I don't think I did anything special other than following <a href=\"https://rwightman.github.io/pytorch-image-models/feature_extraction/\" target=\"_blank\">timm's style</a>.</p>\n<pre><code>import timm\nimport torch\n\n# create convnet\nbackbone = timm.create_model(\n                \"convnext_tiny\",\n                pretrained=True,\n                num_classes=0,\n                global_pool=\"\",\n                in_chans=1,\n                drop_path_rate=0.1,\n            )\nnum_features = backbone.num_features\n\n# inference\nwith torch.inference_mode():\n    mel = torch.rand((1,1,64,500))  # B C H W\n    y_pred = backbone(mel)  # =&gt; (1, 768, 2, 15)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1800967": "Many thanks to Kaggle and Cornell Lab of Ornithology for hosting such an interesting competition.\n\nMy solution is a reimplementation of the BirdCLEF 2021 [2nd place solution](https://www.kaggle.com/competitions/birdclef-2021/discussion/243463).\nSpecial thanks to the new baseline team for publishing the amazing solution. And thanks to @julian3833 for sharing [the great baseline!](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-submit-0-66)\n\n# Model\nMy final submission is an ensemble of 9 models of different seed/fold/backbone, each model using last year's 2nd place solution method.\n\nThe backbones were:\n* 4x `eca_nfnet_l0`\n* 2x `tf_efficientnetv2_s_in21k`\n* 2x `resnet34`\n* 1x `convnext_tiny`\n\n# Oversampling\n\nTo increase the number of files for minority classes(N<20), I split the training files by hand.\nFor example, `maupar` appears only in one file, but the file contains many different types of its songs and calls.\nI cut it into segments of 10-30 seconds using a waveform editor ([Audacity](https://www.audacityteam.org/)) and created multiple training data from a single file.\n\nThe cut audio files were further augmented by applying effects such as noise reduction/reverb/gain/etc. to each segment in the waveform editor. Ultimately, I created 5-20 additional sample files per target minority class.\n\n\n# Training\n\n* Melspec\n\t1. window_size=1024, hop_size=320, fmin=50, fmax=14000, power=2, mel_bins=64, top_db=80 (nfnet, effnet, convnext)\n\t2. window_size=2048, hop_size=512, fmin=16, fmax=16386, power=2, mel_bins=256, top_db=80 (resnet)\n* Data Augmentation(Waveform)\n\t* Background Noise (2020 nocall, 2021 nocall, freefield1010)\n\t* GaussianNoise\n\t* PinkNoise\n\t* NoiseInjection\n\t* RandomVolume\n\t* TimeShift\n* Data Augmentation(Image)\n\t* SpecAug\n\t* CutOut\n\t* Lowpass\n\t* TranslateY (shift in freq dim, simulated pitch shift)\n* Optimizer\n\t* AdamW, LR=1e-3, weight decay=1e-5, Cosine Anearling with warmup\n* Loss\n\t* BCEWithLogitsLoss\n\t* [Weighted loss by rating](https://www.kaggle.com/code/julian3833/birdclef-21-2nd-place-model-train-0-66?scriptVersionId=88985374)\n* 25-30 epochs\n* Training target is the union of primary and secondary labels.\n* All backbones were **pretrained on BirdCLEF2021** data.\n* **micro-f1@0.1** for `scored_birds` was used for validation. Although a bit odd, the correlation between this score and Public/Private LB was not bad.\n\n\n# Post Processing\n\n* Probability averaged over previous and next chunks.\n* Correct thresholds using [call/nocall binary classifier](https://github.com/ChristofHenkel/kaggle-birdclef2021-2nd-place/blob/main/configs/pp_binary_ext3_1.py) probabilities.\n* Threshold optimization by percentile per clesses.\n\n```python\n# thresholds per class\nthreshold = pd.Series(np.percentile(test_df[SCORED_BIRDS].values, 90, axis=0), index=SCORED_BIRDS)\n```\n\n# What Did Not Work\n\n* SED models - could not make it past 0.72 for me\n* CoordConv\n* MC Dropout\n* Label Smoothing\n* Undersampling for majority classes\n* ViT / Swin backbone",
    "1801009": "Congrats on your solo gold and thanks for sharing your solution.",
    "1801653": "Congrats on the solo gold!!",
    "1801669": "Congrats. And you became competition master!",
    "1801708": "Thanks @shinmurashinmura, and congratulations on becoming a Competitions Master too!",
    "1801869": "Thank you. We are competition master.",
    "1802192": "Congrats! Like with 3rd place I am glad that our last year's solution was still working so well here. Thanks for the writeup.",
    "1802260": "Congratulations! 🎉🎉",
    "1802340": "Congrats for this result!\n\nI also tried the convext family but could not train it. I tried resizing the images to 224 x 224 (if I remember it well) in the albumentations step but it still did not solve the size mismatch error in layers. Did you make any pre-process or resizing step to be able to use convnext? I have never used this architecture before. I tried to use it as encoder to subsequent SED model, it did not work...",
    "1802714": "Thanks!\n\nThis is my implementation, but I don't think I did anything special other than following [timm's style](https://rwightman.github.io/pytorch-image-models/feature_extraction/).\n\n```python\nimport timm\nimport torch\n\n# create convnet\nbackbone = timm.create_model(\n                \"convnext_tiny\",\n                pretrained=True,\n                num_classes=0,\n                global_pool=\"\",\n                in_chans=1,\n                drop_path_rate=0.1,\n            )\nnum_features = backbone.num_features\n\n# inference\nwith torch.inference_mode():\n    mel = torch.rand((1,1,64,500))  # B C H W\n    y_pred = backbone(mel)  # => (1, 768, 2, 15)\n```"
  },
  "source": "meta"
}