{
  "id": 412922,
  "title": "top 7th solution - `sumix` augmentation did all the work",
  "url": "/competitions/birdclef-2023/writeups/msu-ysda-hse-top-7th-solution-sumix-augmentation-d",
  "author_name": "",
  "post_date": "2023-05-28T10:21:55.187Z",
  "votes": 38,
  "comment_count": 1,
  "views": 0,
  "content": "<h4>What's important in <a href=\"https://www.kaggle.com/code/loonypenguin/ensemble-openvino-models-inference-notebook\" target=\"_blank\">our best inference notebook</a> ?</h4>\n<ul>\n<li>Ensemble of 19 models</li>\n<li>Ensemble of <code>efficientnet_b2</code> from <code>torchvision</code> and <code>rexnet150</code> from <code>timm</code></li>\n<li>Speedup with <code>openvino</code></li>\n<li>For submission we trained all models with knowledge distillation on full dataset - (<code>rkl</code> prefix is indicator)</li>\n</ul>\n<p>Our five models from 19 with <code>rkl-rx-wdfbn-rx-ema-200-full</code> prefix achieve 0.7471 on private test set and  0.83194 on public test set, <code>rx</code> - <code>rexnet150</code>, distillated <code>rexnet150</code> models from <code>rexnet150</code> trained before</p>\n<h4>Custom Augmentations</h4>\n<p>Let's imagine you're at some party right now and you're surrounded by talking people. It's not hard for you to pay attention to a certain group to hear their conversation and you can freely change your target. So the task - to cut specific speaker from audio - is natural. What we can say about audio in which - one person cries and another one laughs? In such audio there are two actions - laughing and crying, and this is true for birds too. With this consideration we've come up to <code>sumup</code> idea:</p>\n<pre><code>def sumup(waves: torch.Tensor, : torch.Tensor):\n    batch_size = len()\n    perm = torch.randperm(batch_size)\n\n    waves = waves + waves[perm]\n\n     {\n        : waves,\n        : torch.clip( + [perm], =, =)\n    }\n</code></pre>\n<h6>Boosting score</h6>\n<p>What if you have an audio where one bird is more louder than another? This audio is still with two different birds, but if signal from second bird is tiny, it's okay to assume, that maybe probability of second bird is small or even zero. </p>\n<pre><code>def sumix(waves: torch.Tensor, : torch.Tensor, max_percent:  = , min_percent:  = ):\n    batch_size = len()\n    perm = torch.randperm(batch_size)\n    coeffs_1 = torch.rand(batch_size, device=waves.device).(-, ) * (\n        max_percent  - min_percent\n    ) + min_percent\n    coeffs_2 = torch.rand(batch_size, device=waves.device).(-, ) * (\n        max_percent  - min_percent\n    ) + min_percent\n    label_coeffs_1 = torch.where(coeffs_1 &gt;= , ,  -  * ( - coeffs_1))\n    label_coeffs_2 = torch.where(coeffs_2 &gt;= , ,  -  * ( - coeffs_2))\n     = label_coeffs_1 *  + label_coeffs_2 * [perm]\n\n    waves = coeffs_1 * waves + coeffs_2 * waves[perm]\n     {\n        : waves,\n        : torch.clip(, , )\n    }\n</code></pre>\n<p>We propose yet another customization of data mix specific for audio domain - <code>sumix</code> augmentation. As you can see, we used linear slope in our experiments, but it's open question what's slope for labels better to use, maybe <code>min_percent = 0.3</code> doesn't need slope at all. This augmentation is applicable to any audio classification task. <code>sumix</code> is better than <code>sumup</code> because it also affects loudness, i. e. implies another audio augmentation. </p>\n<p>When we decided to check solutions from past competitions, we tried to add background noise in our pipeline, but noise didn't show any improvement. Our thoughts - incorporating new background noises for audios is already in <code>sumix</code> or <code>sumup</code>. For gaussian/colored noises we have similar explanation, why those augmentations didn't work - competition's data is not excellent and consists of noisy samples, so <code>sumix</code> and <code>sumup</code> propagate noise and make noise diverse by themselves.</p>\n<h4>Training pipeline</h4>\n<ul>\n<li><code>concatmix</code> with probability 0.5 - <a href=\"https://github.com/dazzle-me/birdclef-2022-3rd-place-solution/blob/main/experiments/new-metrics-exp-17/fold_4/saved_src/models/layers/concat_mix.py\" target=\"_blank\">implementation</a></li>\n<li><code>sumix(min_percent=0.3, max_percent=1.0)</code> with probability 1</li>\n<li><code>mixup</code> on mel-spectrograms with <code>Beta(1.5, 1.5)</code> with probability 1</li>\n<li><code>cutmix</code> on mel-spectrograms with <code>Beta(1.5, 1.5)</code> with probability 0.5</li>\n<li><code>AdamW</code> with constant <code>LR</code> 1e-3 and <code>weight_decay</code> 1e-2 for all models</li>\n<li><code>binary_cross_entropy</code> for every class, labels are primary labels plus secondary labels</li>\n<li>Random 5 continuous seconds from audio at training, first 5 seconds at validation</li>\n<li>We turned off <code>weight_decay</code> for all normalization layers</li>\n<li><code>EMA</code> with <code>decay</code>  0.999</li>\n<li>Almost total correlation between LB and our CV, ensemble of 20 models gathered as best in CV achieves 0.75021 on private test set and 0.83328 on public test set, this ensemble is our top-3 submission on public test set</li>\n</ul>\n<h4>Knowledge distillation</h4>\n<p>After training a model, we use it as teacher for another model with the same architecture, augmentations, <code>AdamW</code>, <code>EMA</code> and etc., except <code>loss</code>. At second stage with distillation we use $$L = 0.33 * BceFromTargets + 0.34 * KL(probs_{student} || probs_{teacher}) + 0.33 * KL(probs_{teacher} || probs_{student}) $$</p>\n<h4>Models</h4>\n<ul>\n<li><code>rexnet150</code> with <code>dropout</code> 0.3 and <code>drop_path</code> 0.2</li>\n<li><code>rexnet150-time</code> - same <code>rexnet150</code> but with attention head, distilled from <code>rexnet150</code> without <code>time</code>-aware head</li>\n<li><code>efficientnet-b2</code> with <code>dropout</code> 0.3 and <code>drop_path</code> 0.2</li>\n<li><code>efficientnet-b2-time</code> - analogously, distilled from <code>efficientnet-b2</code></li>\n</ul>\n<p>Already mentioned hacks allow to reach score 0.7471 with 5 <code>rexnet150</code>  trained on full data with <code>batch_size=336</code> 200 epochs and knowledge distillation on a single <code>A100 (40GB)</code>. To boost performance we ensemble <code>rexnet150</code>  with <code>efficientnet-b2</code> from <code>torchvision</code>. To boost performance yet another time we use attention head - copy-past <code>self-attention</code> operation from <code>transformers</code>. All experiments were conducted in mix-precision with <code>bfloat16</code> type.</p>\n<p>Ensemble of 5 models with prefix <code>rkl-rx150-wdfbn-rx150-time-ema-ep200-full-{5seeds}</code>, i.e <code>rexnet150</code> with attention head for time awareness distilled from simple <code>rexnet150</code>, achieves 0.75151 on private set and 0.83261 on public set.</p>\n<h4>We didn't try:</h4>\n<ul>\n<li>External data not from BirdClef2021/BirdClef2022</li>\n<li>SSL</li>\n<li>SED models</li>\n<li>Utilization of BirdNet</li>\n<li><code>eca_nfnet</code> from <code>timm</code></li>\n</ul>\n<h4>We tried, but it didn't work:</h4>\n<ul>\n<li>Pretraining our models on BirdClef2021/BirdClef2022, we made few attempts, it didn't show any improvements in LB, experiments conduction became much slower, so we gave up the idea</li>\n<li>Any background noise (even which was used in previous year's competition)</li>\n<li>Any LR scheduler</li>\n<li>Any Noise(Gauss, Colored)</li>\n<li><code>tf_efficientnet_b2_ns</code> from <code>timm</code></li>\n<li><code>SwinV2</code>/<code>ConvNeXt</code> from <code>torchvision</code></li>\n<li>Manifold MixUp</li>\n<li><code>focal_loss</code>, <code>soft_focal_loss</code></li>\n<li>Weighted samplers</li>\n<li><code>SpecAugment</code></li>\n</ul>\n<h4>Speeding up inference time</h4>\n<p>We measured inference time with random input in kaggle kernel, I don't remember which one of intel's processors was, often in our kaggle notebook we got <code>Intel Xeon 2.2 Gz</code>. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13651743%2F0d18919aee04af8f54bfe0880472543d%2Fphoto_2023-05-25%2023.34.07.jpeg?generation=1685046878803412&amp;alt=media\" alt=\"\"></p>\n<p>Our thoughts - <code>ONNX</code> by itself utilizes 2 cpus, <code>OpenVino</code> is faster cause it additionally fuses some layers for inference and uses compilation specific to a certain CPU in submission notebook to take advantage of processor's type</p>\n<h4>Appendix: <code>MultiHeadAttentionClassifier</code>'s implementation. All hyperparameters are preserved</h4>\n<pre><code></code></pre>",
  "messages": [
    {
      "id": "2274329",
      "postDate": "05/25/2023 20:45:53",
      "content": "<h4>What's important in <a href=\"https://www.kaggle.com/code/loonypenguin/ensemble-openvino-models-inference-notebook\" target=\"_blank\">our best inference notebook</a> ?</h4>\n<ul>\n<li>Ensemble of 19 models</li>\n<li>Ensemble of <code>efficientnet_b2</code> from <code>torchvision</code> and <code>rexnet150</code> from <code>timm</code></li>\n<li>Speedup with <code>openvino</code></li>\n<li>For submission we trained all models with knowledge distillation on full dataset - (<code>rkl</code> prefix is indicator)</li>\n</ul>\n<p>Our five models from 19 with <code>rkl-rx-wdfbn-rx-ema-200-full</code> prefix achieve 0.7471 on private test set and  0.83194 on public test set, <code>rx</code> - <code>rexnet150</code>, distillated <code>rexnet150</code> models from <code>rexnet150</code> trained before</p>\n<h4>Custom Augmentations</h4>\n<p>Let's imagine you're at some party right now and you're surrounded by talking people. It's not hard for you to pay attention to a certain group to hear their conversation and you can freely change your target. So the task - to cut specific speaker from audio - is natural. What we can say about audio in which - one person cries and another one laughs? In such audio there are two actions - laughing and crying, and this is true for birds too. With this consideration we've come up to <code>sumup</code> idea:</p>\n<pre><code>def sumup(waves: torch.Tensor, : torch.Tensor):\n    batch_size = len()\n    perm = torch.randperm(batch_size)\n\n    waves = waves + waves[perm]\n\n     {\n        : waves,\n        : torch.clip( + [perm], =, =)\n    }\n</code></pre>\n<h6>Boosting score</h6>\n<p>What if you have an audio where one bird is more louder than another? This audio is still with two different birds, but if signal from second bird is tiny, it's okay to assume, that maybe probability of second bird is small or even zero. </p>\n<pre><code>def sumix(waves: torch.Tensor, : torch.Tensor, max_percent:  = , min_percent:  = ):\n    batch_size = len()\n    perm = torch.randperm(batch_size)\n    coeffs_1 = torch.rand(batch_size, device=waves.device).(-, ) * (\n        max_percent  - min_percent\n    ) + min_percent\n    coeffs_2 = torch.rand(batch_size, device=waves.device).(-, ) * (\n        max_percent  - min_percent\n    ) + min_percent\n    label_coeffs_1 = torch.where(coeffs_1 &gt;= , ,  -  * ( - coeffs_1))\n    label_coeffs_2 = torch.where(coeffs_2 &gt;= , ,  -  * ( - coeffs_2))\n     = label_coeffs_1 *  + label_coeffs_2 * [perm]\n\n    waves = coeffs_1 * waves + coeffs_2 * waves[perm]\n     {\n        : waves,\n        : torch.clip(, , )\n    }\n</code></pre>\n<p>We propose yet another customization of data mix specific for audio domain - <code>sumix</code> augmentation. As you can see, we used linear slope in our experiments, but it's open question what's slope for labels better to use, maybe <code>min_percent = 0.3</code> doesn't need slope at all. This augmentation is applicable to any audio classification task. <code>sumix</code> is better than <code>sumup</code> because it also affects loudness, i. e. implies another audio augmentation. </p>\n<p>When we decided to check solutions from past competitions, we tried to add background noise in our pipeline, but noise didn't show any improvement. Our thoughts - incorporating new background noises for audios is already in <code>sumix</code> or <code>sumup</code>. For gaussian/colored noises we have similar explanation, why those augmentations didn't work - competition's data is not excellent and consists of noisy samples, so <code>sumix</code> and <code>sumup</code> propagate noise and make noise diverse by themselves.</p>\n<h4>Training pipeline</h4>\n<ul>\n<li><code>concatmix</code> with probability 0.5 - <a href=\"https://github.com/dazzle-me/birdclef-2022-3rd-place-solution/blob/main/experiments/new-metrics-exp-17/fold_4/saved_src/models/layers/concat_mix.py\" target=\"_blank\">implementation</a></li>\n<li><code>sumix(min_percent=0.3, max_percent=1.0)</code> with probability 1</li>\n<li><code>mixup</code> on mel-spectrograms with <code>Beta(1.5, 1.5)</code> with probability 1</li>\n<li><code>cutmix</code> on mel-spectrograms with <code>Beta(1.5, 1.5)</code> with probability 0.5</li>\n<li><code>AdamW</code> with constant <code>LR</code> 1e-3 and <code>weight_decay</code> 1e-2 for all models</li>\n<li><code>binary_cross_entropy</code> for every class, labels are primary labels plus secondary labels</li>\n<li>Random 5 continuous seconds from audio at training, first 5 seconds at validation</li>\n<li>We turned off <code>weight_decay</code> for all normalization layers</li>\n<li><code>EMA</code> with <code>decay</code>  0.999</li>\n<li>Almost total correlation between LB and our CV, ensemble of 20 models gathered as best in CV achieves 0.75021 on private test set and 0.83328 on public test set, this ensemble is our top-3 submission on public test set</li>\n</ul>\n<h4>Knowledge distillation</h4>\n<p>After training a model, we use it as teacher for another model with the same architecture, augmentations, <code>AdamW</code>, <code>EMA</code> and etc., except <code>loss</code>. At second stage with distillation we use $$L = 0.33 * BceFromTargets + 0.34 * KL(probs_{student} || probs_{teacher}) + 0.33 * KL(probs_{teacher} || probs_{student}) $$</p>\n<h4>Models</h4>\n<ul>\n<li><code>rexnet150</code> with <code>dropout</code> 0.3 and <code>drop_path</code> 0.2</li>\n<li><code>rexnet150-time</code> - same <code>rexnet150</code> but with attention head, distilled from <code>rexnet150</code> without <code>time</code>-aware head</li>\n<li><code>efficientnet-b2</code> with <code>dropout</code> 0.3 and <code>drop_path</code> 0.2</li>\n<li><code>efficientnet-b2-time</code> - analogously, distilled from <code>efficientnet-b2</code></li>\n</ul>\n<p>Already mentioned hacks allow to reach score 0.7471 with 5 <code>rexnet150</code>  trained on full data with <code>batch_size=336</code> 200 epochs and knowledge distillation on a single <code>A100 (40GB)</code>. To boost performance we ensemble <code>rexnet150</code>  with <code>efficientnet-b2</code> from <code>torchvision</code>. To boost performance yet another time we use attention head - copy-past <code>self-attention</code> operation from <code>transformers</code>. All experiments were conducted in mix-precision with <code>bfloat16</code> type.</p>\n<p>Ensemble of 5 models with prefix <code>rkl-rx150-wdfbn-rx150-time-ema-ep200-full-{5seeds}</code>, i.e <code>rexnet150</code> with attention head for time awareness distilled from simple <code>rexnet150</code>, achieves 0.75151 on private set and 0.83261 on public set.</p>\n<h4>We didn't try:</h4>\n<ul>\n<li>External data not from BirdClef2021/BirdClef2022</li>\n<li>SSL</li>\n<li>SED models</li>\n<li>Utilization of BirdNet</li>\n<li><code>eca_nfnet</code> from <code>timm</code></li>\n</ul>\n<h4>We tried, but it didn't work:</h4>\n<ul>\n<li>Pretraining our models on BirdClef2021/BirdClef2022, we made few attempts, it didn't show any improvements in LB, experiments conduction became much slower, so we gave up the idea</li>\n<li>Any background noise (even which was used in previous year's competition)</li>\n<li>Any LR scheduler</li>\n<li>Any Noise(Gauss, Colored)</li>\n<li><code>tf_efficientnet_b2_ns</code> from <code>timm</code></li>\n<li><code>SwinV2</code>/<code>ConvNeXt</code> from <code>torchvision</code></li>\n<li>Manifold MixUp</li>\n<li><code>focal_loss</code>, <code>soft_focal_loss</code></li>\n<li>Weighted samplers</li>\n<li><code>SpecAugment</code></li>\n</ul>\n<h4>Speeding up inference time</h4>\n<p>We measured inference time with random input in kaggle kernel, I don't remember which one of intel's processors was, often in our kaggle notebook we got <code>Intel Xeon 2.2 Gz</code>. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13651743%2F0d18919aee04af8f54bfe0880472543d%2Fphoto_2023-05-25%2023.34.07.jpeg?generation=1685046878803412&amp;alt=media\" alt=\"\"></p>\n<p>Our thoughts - <code>ONNX</code> by itself utilizes 2 cpus, <code>OpenVino</code> is faster cause it additionally fuses some layers for inference and uses compilation specific to a certain CPU in submission notebook to take advantage of processor's type</p>\n<h4>Appendix: <code>MultiHeadAttentionClassifier</code>'s implementation. All hyperparameters are preserved</h4>\n<pre><code></code></pre>",
      "rawMarkdown": "#### What's important in [our best inference notebook](https://www.kaggle.com/code/loonypenguin/ensemble-openvino-models-inference-notebook) ?\n\n- Ensemble of 19 models\n- Ensemble of `efficientnet_b2` from `torchvision` and `rexnet150` from `timm`\n- Speedup with `openvino`\n- For submission we trained all models with knowledge distillation on full dataset - (`rkl` prefix is indicator)\n\n\nOur five models from 19 with `rkl-rx-wdfbn-rx-ema-200-full` prefix achieve 0.7471 on private test set and  0.83194 on public test set, `rx` - `rexnet150`, distillated `rexnet150` models from `rexnet150` trained before\n\n#### Custom Augmentations\n\nLet's imagine you're at some party right now and you're surrounded by talking people. It's not hard for you to pay attention to a certain group to hear their conversation and you can freely change your target. So the task - to cut specific speaker from audio - is natural. What we can say about audio in which - one person cries and another one laughs? In such audio there are two actions - laughing and crying, and this is true for birds too. With this consideration we've come up to `sumup` idea:\n\n```python3\ndef sumup(waves: torch.Tensor, labels: torch.Tensor):\n    batch_size = len(labels)\n    perm = torch.randperm(batch_size)\n\n    waves = waves + waves[perm]\n\n    return {\n        \"waves\": waves,\n        \"labels\": torch.clip(labels + labels[perm], min=0, max=1)\n    }\n```\n\n###### Boosting score\nWhat if you have an audio where one bird is more louder than another? This audio is still with two different birds, but if signal from second bird is tiny, it's okay to assume, that maybe probability of second bird is small or even zero. \n\n```python3\ndef sumix(waves: torch.Tensor, labels: torch.Tensor, max_percent: float = 1.0, min_percent: float = 0.3):\n    batch_size = len(labels)\n    perm = torch.randperm(batch_size)\n    coeffs_1 = torch.rand(batch_size, device=waves.device).view(-1, 1) * (\n        max_percent  - min_percent\n    ) + min_percent\n    coeffs_2 = torch.rand(batch_size, device=waves.device).view(-1, 1) * (\n        max_percent  - min_percent\n    ) + min_percent\n    label_coeffs_1 = torch.where(coeffs_1 >= 0.5, 1, 1 - 2 * (0.5 - coeffs_1))\n    label_coeffs_2 = torch.where(coeffs_2 >= 0.5, 1, 1 - 2 * (0.5 - coeffs_2))\n    labels = label_coeffs_1 * labels + label_coeffs_2 * labels[perm]\n\n    waves = coeffs_1 * waves + coeffs_2 * waves[perm]\n    return {\n        \"waves\": waves,\n        \"labels\": torch.clip(labels, 0, 1)\n    }\n```\n\nWe propose yet another customization of data mix specific for audio domain - `sumix` augmentation. As you can see, we used linear slope in our experiments, but it's open question what's slope for labels better to use, maybe `min_percent = 0.3` doesn't need slope at all. This augmentation is applicable to any audio classification task. `sumix` is better than `sumup` because it also affects loudness, i. e. implies another audio augmentation. \n\nWhen we decided to check solutions from past competitions, we tried to add background noise in our pipeline, but noise didn't show any improvement. Our thoughts - incorporating new background noises for audios is already in `sumix` or `sumup`. For gaussian/colored noises we have similar explanation, why those augmentations didn't work - competition's data is not excellent and consists of noisy samples, so `sumix` and `sumup` propagate noise and make noise diverse by themselves.\n\n#### Training pipeline\n- `concatmix` with probability 0.5 - [implementation](https://github.com/dazzle-me/birdclef-2022-3rd-place-solution/blob/main/experiments/new-metrics-exp-17/fold_4/saved_src/models/layers/concat_mix.py)\n- `sumix(min_percent=0.3, max_percent=1.0)` with probability 1\n- `mixup` on mel-spectrograms with `Beta(1.5, 1.5)` with probability 1\n- `cutmix` on mel-spectrograms with `Beta(1.5, 1.5)` with probability 0.5\n- `AdamW` with constant `LR` 1e-3 and `weight_decay` 1e-2 for all models\n- `binary_cross_entropy` for every class, labels are primary labels plus secondary labels\n- Random 5 continuous seconds from audio at training, first 5 seconds at validation\n- We turned off `weight_decay` for all normalization layers\n- `EMA` with `decay`  0.999\n- Almost total correlation between LB and our CV, ensemble of 20 models gathered as best in CV achieves 0.75021 on private test set and 0.83328 on public test set, this ensemble is our top-3 submission on public test set\n\n#### Knowledge distillation\n\nAfter training a model, we use it as teacher for another model with the same architecture, augmentations, `AdamW`, `EMA` and etc., except `loss`. At second stage with distillation we use $$L = 0.33 * BceFromTargets + 0.34 * KL(probs_{student} || probs_{teacher}) + 0.33 * KL(probs_{teacher} || probs_{student}) $$\n\n\n#### Models\n- `rexnet150` with `dropout` 0.3 and `drop_path` 0.2\n- `rexnet150-time` - same `rexnet150` but with attention head, distilled from `rexnet150` without `time`-aware head\n- `efficientnet-b2` with `dropout` 0.3 and `drop_path` 0.2\n- `efficientnet-b2-time` - analogously, distilled from `efficientnet-b2`\n\nAlready mentioned hacks allow to reach score 0.7471 with 5 `rexnet150`  trained on full data with `batch_size=336` 200 epochs and knowledge distillation on a single `A100 (40GB)`. To boost performance we ensemble `rexnet150`  with `efficientnet-b2` from `torchvision`. To boost performance yet another time we use attention head - copy-past `self-attention` operation from `transformers`. All experiments were conducted in mix-precision with `bfloat16` type.\n\nEnsemble of 5 models with prefix `rkl-rx150-wdfbn-rx150-time-ema-ep200-full-{5seeds}`, i.e `rexnet150` with attention head for time awareness distilled from simple `rexnet150`, achieves 0.75151 on private set and 0.83261 on public set.\n\n#### We didn't try:\n- External data not from BirdClef2021/BirdClef2022\n- SSL\n- SED models\n- Utilization of BirdNet\n- `eca_nfnet` from `timm`\n\n#### We tried, but it didn't work:\n- Pretraining our models on BirdClef2021/BirdClef2022, we made few attempts, it didn't show any improvements in LB, experiments conduction became much slower, so we gave up the idea\n- Any background noise (even which was used in previous year's competition)\n- Any LR scheduler\n- Any Noise(Gauss, Colored)\n- `tf_efficientnet_b2_ns` from `timm`\n- `SwinV2`/`ConvNeXt` from `torchvision`\n- Manifold MixUp\n- `focal_loss`, `soft_focal_loss`\n- Weighted samplers\n- `SpecAugment`\n\n#### Speeding up inference time\n\nWe measured inference time with random input in kaggle kernel, I don't remember which one of intel's processors was, often in our kaggle notebook we got `Intel Xeon 2.2 Gz`. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13651743%2F0d18919aee04af8f54bfe0880472543d%2Fphoto_2023-05-25%2023.34.07.jpeg?generation=1685046878803412&alt=media)\n\nOur thoughts - `ONNX` by itself utilizes 2 cpus, `OpenVino` is faster cause it additionally fuses some layers for inference and uses compilation specific to a certain CPU in submission notebook to take advantage of processor's type\n\n\n#### Appendix: `MultiHeadAttentionClassifier`'s implementation. All hyperparameters are preserved\n```python3\nclass MultiHeadSelfAttention(nn.Module):\n    def __init__(\n        self,\n        input_channel: int,\n        head_size: int,\n        num_heads: int,\n        attention_dropout: float,\n    ) -> None:\n        super().__init__()\n        hidden_dim = head_size * num_heads\n        self.hidden_dim = hidden_dim\n\n        self.head_size = head_size\n        self.num_heads = num_heads\n\n        self.key = nn.Linear(input_channel, hidden_dim)\n        self.query = nn.Linear(input_channel, hidden_dim)\n\n        self.attention_dropout = nn.Dropout(attention_dropout)\n        self.sqrt_head_size = sqrt(head_size)\n\n        self.value = nn.Linear(input_channel, hidden_dim)\n\n    def tranpose_for_scores(self, x: Tensor) -> Tensor:\n        # [BS; T; H] -> [BS; T; K, M]\n        new_x_shape = x.size()[:-1] + (self.num_heads, self.head_size)\n        x = x.view(new_x_shape)\n        return x.permute(0, 2, 1, 3) # [BS; K; T; M]\n\n    def get_key(self, x: Tensor) -> Tensor:\n        return self.tranpose_for_scores(self.key(x))\n\n    def get_query(self, x: Tensor) -> Tensor:\n        return self.tranpose_for_scores(self.query(x))\n\n    def get_value(self, x: Tensor) -> Tensor:\n        return self.tranpose_for_scores(self.value(x))\n\n    def forward(self, x: Tensor) -> Tensor:\n        # [BS; L; C]\n        key = self.get_key(x)\n        query = self.get_query(x)\n        value = self.get_value(x)\n\n        attention_scores = torch.matmul(query, key.transpose(-1, -2)) \n        attention_scores /= self.sqrt_head_size\n        attention_scores = F.softmax(attention_scores, dim=-1)\n        attention_scores = self.attention_dropout(attention_scores)\n        x = torch.matmul(attention_scores, value) # [BS; K; N; M]\n        x = x.permute(0, 2, 1, 3).contiguous()\n        x = x.view(x.shape[:2] + (self.hidden_dim,))\n        return x\n\nclass MultiHeadAttentionClassifier(MultiHeadSelfAttention):\n    def __init__(\n        self,\n        input_channel: int,\n        head_size: int = 32,\n        num_heads: int = 24,\n        attention_dropout: float = 0.3,\n        num_classes: int = 264,\n        dropout: float = 0.3\n    ) -> None:\n        super().__init__(input_channel, head_size, num_heads, attention_dropout)\n        self.query = nn.Parameter(\n            torch.empty(num_heads, num_classes, head_size),\n            requires_grad=True\n        )\n        nn.init.normal_(self.query)\n        self.classifier = nn.Sequential(\n            nn.Dropout(dropout),\n            nn.Conv1d(num_classes, num_classes, kernel_size=self.hidden_dim, groups=num_classes)\n        )\n\n    def get_query(self, x: Tensor) -> Tensor:\n        return self.query\n\n    def forward(self, x: Tensor) -> Tensor:\n        # x.shape == [BS; C; F; T] // output from backbone\n        x = vec.view(x.shape[:-2] + (-1,))\n        # [BS; C; L]\n        x = vec.permute(0, 2, 1)\n       # [BS; L; C]\n        \n        out = super().forward(x)\n        return self.classifier(out).squeeze()\n```",
      "votes": null
    },
    {
      "id": "2286252",
      "postDate": "06/03/2023 10:31:07",
      "content": "<p>Congrats for landing in the gold zone and thanks for sharing the writeup</p>",
      "rawMarkdown": "Congrats for landing in the gold zone and thanks for sharing the writeup",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2286252,
      "author_name": "pardeep19singh",
      "author_url": "",
      "post_date": "06/03/2023 10:31:07",
      "content": "<p>Congrats for landing in the gold zone and thanks for sharing the writeup</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2274329": "#### What's important in [our best inference notebook](https://www.kaggle.com/code/loonypenguin/ensemble-openvino-models-inference-notebook) ?\n\n- Ensemble of 19 models\n- Ensemble of `efficientnet_b2` from `torchvision` and `rexnet150` from `timm`\n- Speedup with `openvino`\n- For submission we trained all models with knowledge distillation on full dataset - (`rkl` prefix is indicator)\n\n\nOur five models from 19 with `rkl-rx-wdfbn-rx-ema-200-full` prefix achieve 0.7471 on private test set and  0.83194 on public test set, `rx` - `rexnet150`, distillated `rexnet150` models from `rexnet150` trained before\n\n#### Custom Augmentations\n\nLet's imagine you're at some party right now and you're surrounded by talking people. It's not hard for you to pay attention to a certain group to hear their conversation and you can freely change your target. So the task - to cut specific speaker from audio - is natural. What we can say about audio in which - one person cries and another one laughs? In such audio there are two actions - laughing and crying, and this is true for birds too. With this consideration we've come up to `sumup` idea:\n\n```python3\ndef sumup(waves: torch.Tensor, labels: torch.Tensor):\n    batch_size = len(labels)\n    perm = torch.randperm(batch_size)\n\n    waves = waves + waves[perm]\n\n    return {\n        \"waves\": waves,\n        \"labels\": torch.clip(labels + labels[perm], min=0, max=1)\n    }\n```\n\n###### Boosting score\nWhat if you have an audio where one bird is more louder than another? This audio is still with two different birds, but if signal from second bird is tiny, it's okay to assume, that maybe probability of second bird is small or even zero. \n\n```python3\ndef sumix(waves: torch.Tensor, labels: torch.Tensor, max_percent: float = 1.0, min_percent: float = 0.3):\n    batch_size = len(labels)\n    perm = torch.randperm(batch_size)\n    coeffs_1 = torch.rand(batch_size, device=waves.device).view(-1, 1) * (\n        max_percent  - min_percent\n    ) + min_percent\n    coeffs_2 = torch.rand(batch_size, device=waves.device).view(-1, 1) * (\n        max_percent  - min_percent\n    ) + min_percent\n    label_coeffs_1 = torch.where(coeffs_1 >= 0.5, 1, 1 - 2 * (0.5 - coeffs_1))\n    label_coeffs_2 = torch.where(coeffs_2 >= 0.5, 1, 1 - 2 * (0.5 - coeffs_2))\n    labels = label_coeffs_1 * labels + label_coeffs_2 * labels[perm]\n\n    waves = coeffs_1 * waves + coeffs_2 * waves[perm]\n    return {\n        \"waves\": waves,\n        \"labels\": torch.clip(labels, 0, 1)\n    }\n```\n\nWe propose yet another customization of data mix specific for audio domain - `sumix` augmentation. As you can see, we used linear slope in our experiments, but it's open question what's slope for labels better to use, maybe `min_percent = 0.3` doesn't need slope at all. This augmentation is applicable to any audio classification task. `sumix` is better than `sumup` because it also affects loudness, i. e. implies another audio augmentation. \n\nWhen we decided to check solutions from past competitions, we tried to add background noise in our pipeline, but noise didn't show any improvement. Our thoughts - incorporating new background noises for audios is already in `sumix` or `sumup`. For gaussian/colored noises we have similar explanation, why those augmentations didn't work - competition's data is not excellent and consists of noisy samples, so `sumix` and `sumup` propagate noise and make noise diverse by themselves.\n\n#### Training pipeline\n- `concatmix` with probability 0.5 - [implementation](https://github.com/dazzle-me/birdclef-2022-3rd-place-solution/blob/main/experiments/new-metrics-exp-17/fold_4/saved_src/models/layers/concat_mix.py)\n- `sumix(min_percent=0.3, max_percent=1.0)` with probability 1\n- `mixup` on mel-spectrograms with `Beta(1.5, 1.5)` with probability 1\n- `cutmix` on mel-spectrograms with `Beta(1.5, 1.5)` with probability 0.5\n- `AdamW` with constant `LR` 1e-3 and `weight_decay` 1e-2 for all models\n- `binary_cross_entropy` for every class, labels are primary labels plus secondary labels\n- Random 5 continuous seconds from audio at training, first 5 seconds at validation\n- We turned off `weight_decay` for all normalization layers\n- `EMA` with `decay`  0.999\n- Almost total correlation between LB and our CV, ensemble of 20 models gathered as best in CV achieves 0.75021 on private test set and 0.83328 on public test set, this ensemble is our top-3 submission on public test set\n\n#### Knowledge distillation\n\nAfter training a model, we use it as teacher for another model with the same architecture, augmentations, `AdamW`, `EMA` and etc., except `loss`. At second stage with distillation we use $$L = 0.33 * BceFromTargets + 0.34 * KL(probs_{student} || probs_{teacher}) + 0.33 * KL(probs_{teacher} || probs_{student}) $$\n\n\n#### Models\n- `rexnet150` with `dropout` 0.3 and `drop_path` 0.2\n- `rexnet150-time` - same `rexnet150` but with attention head, distilled from `rexnet150` without `time`-aware head\n- `efficientnet-b2` with `dropout` 0.3 and `drop_path` 0.2\n- `efficientnet-b2-time` - analogously, distilled from `efficientnet-b2`\n\nAlready mentioned hacks allow to reach score 0.7471 with 5 `rexnet150`  trained on full data with `batch_size=336` 200 epochs and knowledge distillation on a single `A100 (40GB)`. To boost performance we ensemble `rexnet150`  with `efficientnet-b2` from `torchvision`. To boost performance yet another time we use attention head - copy-past `self-attention` operation from `transformers`. All experiments were conducted in mix-precision with `bfloat16` type.\n\nEnsemble of 5 models with prefix `rkl-rx150-wdfbn-rx150-time-ema-ep200-full-{5seeds}`, i.e `rexnet150` with attention head for time awareness distilled from simple `rexnet150`, achieves 0.75151 on private set and 0.83261 on public set.\n\n#### We didn't try:\n- External data not from BirdClef2021/BirdClef2022\n- SSL\n- SED models\n- Utilization of BirdNet\n- `eca_nfnet` from `timm`\n\n#### We tried, but it didn't work:\n- Pretraining our models on BirdClef2021/BirdClef2022, we made few attempts, it didn't show any improvements in LB, experiments conduction became much slower, so we gave up the idea\n- Any background noise (even which was used in previous year's competition)\n- Any LR scheduler\n- Any Noise(Gauss, Colored)\n- `tf_efficientnet_b2_ns` from `timm`\n- `SwinV2`/`ConvNeXt` from `torchvision`\n- Manifold MixUp\n- `focal_loss`, `soft_focal_loss`\n- Weighted samplers\n- `SpecAugment`\n\n#### Speeding up inference time\n\nWe measured inference time with random input in kaggle kernel, I don't remember which one of intel's processors was, often in our kaggle notebook we got `Intel Xeon 2.2 Gz`. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13651743%2F0d18919aee04af8f54bfe0880472543d%2Fphoto_2023-05-25%2023.34.07.jpeg?generation=1685046878803412&alt=media)\n\nOur thoughts - `ONNX` by itself utilizes 2 cpus, `OpenVino` is faster cause it additionally fuses some layers for inference and uses compilation specific to a certain CPU in submission notebook to take advantage of processor's type\n\n\n#### Appendix: `MultiHeadAttentionClassifier`'s implementation. All hyperparameters are preserved\n```python3\nclass MultiHeadSelfAttention(nn.Module):\n    def __init__(\n        self,\n        input_channel: int,\n        head_size: int,\n        num_heads: int,\n        attention_dropout: float,\n    ) -> None:\n        super().__init__()\n        hidden_dim = head_size * num_heads\n        self.hidden_dim = hidden_dim\n\n        self.head_size = head_size\n        self.num_heads = num_heads\n\n        self.key = nn.Linear(input_channel, hidden_dim)\n        self.query = nn.Linear(input_channel, hidden_dim)\n\n        self.attention_dropout = nn.Dropout(attention_dropout)\n        self.sqrt_head_size = sqrt(head_size)\n\n        self.value = nn.Linear(input_channel, hidden_dim)\n\n    def tranpose_for_scores(self, x: Tensor) -> Tensor:\n        # [BS; T; H] -> [BS; T; K, M]\n        new_x_shape = x.size()[:-1] + (self.num_heads, self.head_size)\n        x = x.view(new_x_shape)\n        return x.permute(0, 2, 1, 3) # [BS; K; T; M]\n\n    def get_key(self, x: Tensor) -> Tensor:\n        return self.tranpose_for_scores(self.key(x))\n\n    def get_query(self, x: Tensor) -> Tensor:\n        return self.tranpose_for_scores(self.query(x))\n\n    def get_value(self, x: Tensor) -> Tensor:\n        return self.tranpose_for_scores(self.value(x))\n\n    def forward(self, x: Tensor) -> Tensor:\n        # [BS; L; C]\n        key = self.get_key(x)\n        query = self.get_query(x)\n        value = self.get_value(x)\n\n        attention_scores = torch.matmul(query, key.transpose(-1, -2)) \n        attention_scores /= self.sqrt_head_size\n        attention_scores = F.softmax(attention_scores, dim=-1)\n        attention_scores = self.attention_dropout(attention_scores)\n        x = torch.matmul(attention_scores, value) # [BS; K; N; M]\n        x = x.permute(0, 2, 1, 3).contiguous()\n        x = x.view(x.shape[:2] + (self.hidden_dim,))\n        return x\n\nclass MultiHeadAttentionClassifier(MultiHeadSelfAttention):\n    def __init__(\n        self,\n        input_channel: int,\n        head_size: int = 32,\n        num_heads: int = 24,\n        attention_dropout: float = 0.3,\n        num_classes: int = 264,\n        dropout: float = 0.3\n    ) -> None:\n        super().__init__(input_channel, head_size, num_heads, attention_dropout)\n        self.query = nn.Parameter(\n            torch.empty(num_heads, num_classes, head_size),\n            requires_grad=True\n        )\n        nn.init.normal_(self.query)\n        self.classifier = nn.Sequential(\n            nn.Dropout(dropout),\n            nn.Conv1d(num_classes, num_classes, kernel_size=self.hidden_dim, groups=num_classes)\n        )\n\n    def get_query(self, x: Tensor) -> Tensor:\n        return self.query\n\n    def forward(self, x: Tensor) -> Tensor:\n        # x.shape == [BS; C; F; T] // output from backbone\n        x = vec.view(x.shape[:-2] + (-1,))\n        # [BS; C; L]\n        x = vec.permute(0, 2, 1)\n       # [BS; L; C]\n        \n        out = super().forward(x)\n        return self.classifier(out).squeeze()\n```",
    "2286252": "Congrats for landing in the gold zone and thanks for sharing the writeup"
  },
  "source": "meta"
}