{
  "id": 327019,
  "title": "8th place solution. And gold on last attempt",
  "url": "/competitions/birdclef-2022/discussion/327019",
  "author_name": "Kolya Forrat",
  "post_date": "2022-05-25T09:28:36.956000",
  "votes": 32,
  "comment_count": 8,
  "views": 0,
  "content": "<p>It was my first Kaggle competition on which I spent a lot of time. I'm honor to be part of it, big thanks to hosts, Kaggle and all participants.</p>\n<h2>First step</h2>\n<ol>\n<li><p>I started with <a href=\"https://www.kaggle.com/kaerunantoka\" target=\"_blank\">@kaerunantoka</a> public <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">notebook</a> with change mel spectrogram hop length and input shape [224, 512] and add <strong>class weights</strong> to a submission.<br>\nWeights array is 500 divided to amount of each birds species, and clamp it to max value 10. Then I multiply model output to weights array. <strong>0.77</strong> Public LB.</p></li>\n<li><p>I tried this approach with secondary labels with 0.3, 0.4 and 0.5 labels. And with ensembling I'v got <strong>0.79</strong> on LB</p></li>\n</ol>\n<h5>Augmentations</h5>\n<p>For waveform:</p>\n<pre><code>Compose([OneOf(\n     [AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.011, p=1),\n            NoiseInjection(p=1, max_noise_level=0.04)], p=0.4),                                \n            PitchShift(min_semitones=-4, max_semitones=4, p=0.1),\n            Shift(min_fraction=-0.5, max_fraction=0.5, p=0.1),\n            Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),    \n            Normalize(p=1.) ])\n</code></pre>\n<p>For spectrogram:</p>\n<pre><code>torchaudio.transforms.FrequencyMasking(24)\ntorchaudio.transforms.TimeMasking(96)\n</code></pre>\n<h5>Model</h5>\n<p>SED with tf_efficientnet_b0_ns backbone</p>\n<h5>Training</h5>\n<ul>\n<li>Epochs = 30</li>\n<li>LR = 0.001</li>\n<li>weight_decay = 0.0001</li>\n<li>dropout = 0.4</li>\n<li>Loss: Focal BCE Loss</li>\n<li>Optimizer: Adam (betas=(0.95, 0.999))</li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<h5>Validation</h5>\n<p>I validate models on 7 folds CV with f1-macro metric. Thresholds for f1: [0.5, 0.1, 0.15, 0.20, 0.25, 0.30, 0.35, 0.4, 0.45, 0.5]. Best checkpoints chosen by 0.3 and 0.5 thresholds.</p>\n<h5>What didn't work on this step</h5>\n<ol>\n<li>Noise reduction</li>\n<li>Oversampling rare birds</li>\n<li>Trim long silence with noise reduction</li>\n<li>Weighted loss</li>\n<li>3 channels input, where 1st channel is power_to_db, 2nd melspectrogram, 3rd normalized melspectrogram</li>\n</ol>\n<h2>Second step</h2>\n<ol>\n<li>I manually trim all segments without bird's sounds from the scored audios, except skylar and houfin, on them I processed only 70 records per class. </li>\n<li>Splited data to 15 seconds chunks for audios with length less than 1 minute and 30 seconds chunks for more than 1 minute. Got ~40.000 records.</li>\n<li>Used weights arrays with 0.25-0.75 power to reduce their impact to inference.</li>\n<li>Training on this data with previous pipeline and ensemble with previous models gives me <strong>0.81</strong> on LB</li>\n</ol>\n<h5>Data preprocessing</h5>\n<p>Same</p>\n<h5>Augmentations</h5>\n<p>Same</p>\n<h5>Model</h5>\n<p>SED with tf_efficientnet_b0_ns backbone</p>\n<h5>Training</h5>\n<ul>\n<li>Epochs = 60</li>\n<li>LR = 0.0013</li>\n<li>weight_decay = 0.0001</li>\n<li>dropout = 0.4</li>\n<li>Loss: Focal BCE Loss</li>\n<li>Optimizer: Adam (betas=(0.95, 0.999))</li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<h5>What didn't work on this step</h5>\n<ol>\n<li>PaSST model</li>\n<li>AST model</li>\n<li>Linear head models</li>\n<li>PaSST preprocessing</li>\n</ol>\n<h2>Third step</h2>\n<ol>\n<li>I used similar with AST preprocessing.</li>\n<li>Every epoch from all data randomly choose up to 300 records for every class.</li>\n<li>Train on random 5 seconds crop, validate on first 5 seconds.</li>\n<li>Secondary labels 0.4</li>\n<li>Class weights array clamped with max value 8. And used with power 0.6. </li>\n<li>Used mixup 0.4 for first 15 epochs, mixup 0.07 for 16-22 epochs, and no mixup for others.</li>\n<li>Ensembling of this approach with previous gives me <strong>0.82</strong> LB</li>\n</ol>\n<h5>Data preprocessing</h5>\n<pre><code>waveform, sr = ta.load(filename)\nwaveform = crop_or_pad(waveform, sr=SR, mode=self.mode)\nwaveform = waveform - waveform.mean()\nwaveform = torch.tensor(self.wave_transforms(samples=waveform[0].numpy(), sample_rate=SR)).unsqueeze(0)\nfbank = ta.compliance.kaldi.fbank(waveform, htk_compat=True,          sample_frequency=SR, use_energy=False, window_type='hanning',                 num_mel_bins=self.melbins, dither=0.0, frame_shift=9.7)\nfbank = (fbank - self.norm_mean) / (self.norm_std * 2)\n</code></pre>\n<p>Mean and std calculated for all train dataset.</p>\n<h5>Model</h5>\n<p>SED with tf_efficientnet_b0_ns backbone</p>\n<h5>Augmentations</h5>\n<p>Same</p>\n<h5>Training</h5>\n<ul>\n<li>Epochs = 85</li>\n<li>mixup 0.4, when epochs &lt; 15</li>\n<li>LR = 0.0008</li>\n<li>weight_decay = 0.0001</li>\n<li>dropout = 0.4</li>\n<li>Loss: Focal BCE Loss</li>\n<li>Optimizer: Adam (betas=(0.95, 0.999))</li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<h5>Postprocessing</h5>\n<ul>\n<li>Mean-median averaging of predictions</li>\n</ul>\n<pre><code>          full_med = np.median(full_events, axis=1)                    \n          full_mean = np.mean(full_events, axis=1)\n          full_events = np.mean(np.array([full_med, full_mean]), axis=0) \n</code></pre>\n<ul>\n<li>Max adder.</li>\n</ul>\n<pre><code>       logits_max = full_events.max(0)\n       for jk in range(full_events.shape[1]):\n           if logits_max[jk] &gt; threshold * 2.5:\n               full_events[:, jk] += threshold * 0.5\n</code></pre>\n<h2>Story of gold medal</h2>\n<p>When there were 2 days left until the end of the competition, I thought about using a pseudo labeling. And separate my data by 0.8 threshold to remove all noisy records. Only ~15.000 left from ~40.000. Then I trained a new model on this data. Because there were not enough time I trained model on all train data and validate it on some random part and random part of only scored birds. <br>\nAll training and preprocessing parameters was the same with previous.</p>\n<p>It was last day of competition and only 5 submissions. So the first 3 I lost because i forgot to make model.eval() :)<br>\n4th attempt I used with single new model and get <strong>0.79</strong> LB.<br>\nAnd the last one was mean between my best previous attempt and this pseudo labeling new model.</p>\n<p>And it gave me <strong>0.79</strong> private, when my previous best private was <strong>0.78</strong> </p>",
  "messages": [
    {
      "id": 1800868,
      "postDate": "2022-05-25T09:28:36.957Z",
      "content": "<p>It was my first Kaggle competition on which I spent a lot of time. I'm honor to be part of it, big thanks to hosts, Kaggle and all participants.</p>\n<h2>First step</h2>\n<ol>\n<li><p>I started with <a href=\"https://www.kaggle.com/kaerunantoka\" target=\"_blank\">@kaerunantoka</a> public <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0\" target=\"_blank\">notebook</a> with change mel spectrogram hop length and input shape [224, 512] and add <strong>class weights</strong> to a submission.<br>\nWeights array is 500 divided to amount of each birds species, and clamp it to max value 10. Then I multiply model output to weights array. <strong>0.77</strong> Public LB.</p></li>\n<li><p>I tried this approach with secondary labels with 0.3, 0.4 and 0.5 labels. And with ensembling I'v got <strong>0.79</strong> on LB</p></li>\n</ol>\n<h5>Augmentations</h5>\n<p>For waveform:</p>\n<pre><code>Compose([OneOf(\n     [AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.011, p=1),\n            NoiseInjection(p=1, max_noise_level=0.04)], p=0.4),                                \n            PitchShift(min_semitones=-4, max_semitones=4, p=0.1),\n            Shift(min_fraction=-0.5, max_fraction=0.5, p=0.1),\n            Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),    \n            Normalize(p=1.) ])\n</code></pre>\n<p>For spectrogram:</p>\n<pre><code>torchaudio.transforms.FrequencyMasking(24)\ntorchaudio.transforms.TimeMasking(96)\n</code></pre>\n<h5>Model</h5>\n<p>SED with tf_efficientnet_b0_ns backbone</p>\n<h5>Training</h5>\n<ul>\n<li>Epochs = 30</li>\n<li>LR = 0.001</li>\n<li>weight_decay = 0.0001</li>\n<li>dropout = 0.4</li>\n<li>Loss: Focal BCE Loss</li>\n<li>Optimizer: Adam (betas=(0.95, 0.999))</li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<h5>Validation</h5>\n<p>I validate models on 7 folds CV with f1-macro metric. Thresholds for f1: [0.5, 0.1, 0.15, 0.20, 0.25, 0.30, 0.35, 0.4, 0.45, 0.5]. Best checkpoints chosen by 0.3 and 0.5 thresholds.</p>\n<h5>What didn't work on this step</h5>\n<ol>\n<li>Noise reduction</li>\n<li>Oversampling rare birds</li>\n<li>Trim long silence with noise reduction</li>\n<li>Weighted loss</li>\n<li>3 channels input, where 1st channel is power_to_db, 2nd melspectrogram, 3rd normalized melspectrogram</li>\n</ol>\n<h2>Second step</h2>\n<ol>\n<li>I manually trim all segments without bird's sounds from the scored audios, except skylar and houfin, on them I processed only 70 records per class. </li>\n<li>Splited data to 15 seconds chunks for audios with length less than 1 minute and 30 seconds chunks for more than 1 minute. Got ~40.000 records.</li>\n<li>Used weights arrays with 0.25-0.75 power to reduce their impact to inference.</li>\n<li>Training on this data with previous pipeline and ensemble with previous models gives me <strong>0.81</strong> on LB</li>\n</ol>\n<h5>Data preprocessing</h5>\n<p>Same</p>\n<h5>Augmentations</h5>\n<p>Same</p>\n<h5>Model</h5>\n<p>SED with tf_efficientnet_b0_ns backbone</p>\n<h5>Training</h5>\n<ul>\n<li>Epochs = 60</li>\n<li>LR = 0.0013</li>\n<li>weight_decay = 0.0001</li>\n<li>dropout = 0.4</li>\n<li>Loss: Focal BCE Loss</li>\n<li>Optimizer: Adam (betas=(0.95, 0.999))</li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<h5>What didn't work on this step</h5>\n<ol>\n<li>PaSST model</li>\n<li>AST model</li>\n<li>Linear head models</li>\n<li>PaSST preprocessing</li>\n</ol>\n<h2>Third step</h2>\n<ol>\n<li>I used similar with AST preprocessing.</li>\n<li>Every epoch from all data randomly choose up to 300 records for every class.</li>\n<li>Train on random 5 seconds crop, validate on first 5 seconds.</li>\n<li>Secondary labels 0.4</li>\n<li>Class weights array clamped with max value 8. And used with power 0.6. </li>\n<li>Used mixup 0.4 for first 15 epochs, mixup 0.07 for 16-22 epochs, and no mixup for others.</li>\n<li>Ensembling of this approach with previous gives me <strong>0.82</strong> LB</li>\n</ol>\n<h5>Data preprocessing</h5>\n<pre><code>waveform, sr = ta.load(filename)\nwaveform = crop_or_pad(waveform, sr=SR, mode=self.mode)\nwaveform = waveform - waveform.mean()\nwaveform = torch.tensor(self.wave_transforms(samples=waveform[0].numpy(), sample_rate=SR)).unsqueeze(0)\nfbank = ta.compliance.kaldi.fbank(waveform, htk_compat=True,          sample_frequency=SR, use_energy=False, window_type='hanning',                 num_mel_bins=self.melbins, dither=0.0, frame_shift=9.7)\nfbank = (fbank - self.norm_mean) / (self.norm_std * 2)\n</code></pre>\n<p>Mean and std calculated for all train dataset.</p>\n<h5>Model</h5>\n<p>SED with tf_efficientnet_b0_ns backbone</p>\n<h5>Augmentations</h5>\n<p>Same</p>\n<h5>Training</h5>\n<ul>\n<li>Epochs = 85</li>\n<li>mixup 0.4, when epochs &lt; 15</li>\n<li>LR = 0.0008</li>\n<li>weight_decay = 0.0001</li>\n<li>dropout = 0.4</li>\n<li>Loss: Focal BCE Loss</li>\n<li>Optimizer: Adam (betas=(0.95, 0.999))</li>\n<li>Scheduler: CosineAnnealingLR</li>\n</ul>\n<h5>Postprocessing</h5>\n<ul>\n<li>Mean-median averaging of predictions</li>\n</ul>\n<pre><code>          full_med = np.median(full_events, axis=1)                    \n          full_mean = np.mean(full_events, axis=1)\n          full_events = np.mean(np.array([full_med, full_mean]), axis=0) \n</code></pre>\n<ul>\n<li>Max adder.</li>\n</ul>\n<pre><code>       logits_max = full_events.max(0)\n       for jk in range(full_events.shape[1]):\n           if logits_max[jk] &gt; threshold * 2.5:\n               full_events[:, jk] += threshold * 0.5\n</code></pre>\n<h2>Story of gold medal</h2>\n<p>When there were 2 days left until the end of the competition, I thought about using a pseudo labeling. And separate my data by 0.8 threshold to remove all noisy records. Only ~15.000 left from ~40.000. Then I trained a new model on this data. Because there were not enough time I trained model on all train data and validate it on some random part and random part of only scored birds. <br>\nAll training and preprocessing parameters was the same with previous.</p>\n<p>It was last day of competition and only 5 submissions. So the first 3 I lost because i forgot to make model.eval() :)<br>\n4th attempt I used with single new model and get <strong>0.79</strong> LB.<br>\nAnd the last one was mean between my best previous attempt and this pseudo labeling new model.</p>\n<p>And it gave me <strong>0.79</strong> private, when my previous best private was <strong>0.78</strong> </p>",
      "rawMarkdown": "It was my first Kaggle competition on which I spent a lot of time. I'm honor to be part of it, big thanks to hosts, Kaggle and all participants.\n\n## First step\n\n1. I started with @kaerunantoka public [notebook](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0) with change mel spectrogram hop length and input shape [224, 512] and add **class weights** to a submission.\nWeights array is 500 divided to amount of each birds species, and clamp it to max value 10. Then I multiply model output to weights array. **0.77** Public LB.\n\n2. I tried this approach with secondary labels with 0.3, 0.4 and 0.5 labels. And with ensembling I'v got **0.79** on LB\n\n##### Augmentations\nFor waveform:\n```\nCompose([OneOf(\n     [AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.011, p=1),\n            NoiseInjection(p=1, max_noise_level=0.04)], p=0.4),                                \n            PitchShift(min_semitones=-4, max_semitones=4, p=0.1),\n            Shift(min_fraction=-0.5, max_fraction=0.5, p=0.1),\n            Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),    \n            Normalize(p=1.) ])\n```\nFor spectrogram:\n ```\ntorchaudio.transforms.FrequencyMasking(24)\ntorchaudio.transforms.TimeMasking(96)\n```\n##### Model \nSED with tf_efficientnet_b0_ns backbone\n\n##### Training\n* Epochs = 30\n* LR = 0.001\n* weight_decay = 0.0001\n* dropout = 0.4\n* Loss: Focal BCE Loss\n* Optimizer: Adam (betas=(0.95, 0.999))\n* Scheduler: CosineAnnealingLR\n\n##### Validation\nI validate models on 7 folds CV with f1-macro metric. Thresholds for f1: [0.5, 0.1, 0.15, 0.20, 0.25, 0.30, 0.35, 0.4, 0.45, 0.5]. Best checkpoints chosen by 0.3 and 0.5 thresholds.\n\n##### What didn't work on this step\n1. Noise reduction\n2. Oversampling rare birds\n3. Trim long silence with noise reduction\n4. Weighted loss\n5. 3 channels input, where 1st channel is power_to_db, 2nd melspectrogram, 3rd normalized melspectrogram\n\n\n## Second step\n1. I manually trim all segments without bird's sounds from the scored audios, except skylar and houfin, on them I processed only 70 records per class. \n2. Splited data to 15 seconds chunks for audios with length less than 1 minute and 30 seconds chunks for more than 1 minute. Got ~40.000 records.\n3. Used weights arrays with 0.25-0.75 power to reduce their impact to inference.\n4. Training on this data with previous pipeline and ensemble with previous models gives me **0.81** on LB\n\n##### Data preprocessing\nSame\n\n##### Augmentations\nSame\n\n##### Model \nSED with tf_efficientnet_b0_ns backbone\n\n##### Training\n* Epochs = 60\n* LR = 0.0013\n* weight_decay = 0.0001\n* dropout = 0.4\n* Loss: Focal BCE Loss\n* Optimizer: Adam (betas=(0.95, 0.999))\n* Scheduler: CosineAnnealingLR\n\n##### What didn't work on this step\n1. PaSST model\n2. AST model\n3. Linear head models\n4. PaSST preprocessing\n\n## Third step\n1. I used similar with AST preprocessing.\n2. Every epoch from all data randomly choose up to 300 records for every class.\n3. Train on random 5 seconds crop, validate on first 5 seconds.\n4. Secondary labels 0.4\n5. Class weights array clamped with max value 8. And used with power 0.6. \n6. Used mixup 0.4 for first 15 epochs, mixup 0.07 for 16-22 epochs, and no mixup for others.\n7. Ensembling of this approach with previous gives me **0.82** LB\n\n##### Data preprocessing\n```\nwaveform, sr = ta.load(filename)\nwaveform = crop_or_pad(waveform, sr=SR, mode=self.mode)\nwaveform = waveform - waveform.mean()\nwaveform = torch.tensor(self.wave_transforms(samples=waveform[0].numpy(), sample_rate=SR)).unsqueeze(0)\nfbank = ta.compliance.kaldi.fbank(waveform, htk_compat=True,          sample_frequency=SR, use_energy=False, window_type='hanning',                 num_mel_bins=self.melbins, dither=0.0, frame_shift=9.7)\nfbank = (fbank - self.norm_mean) / (self.norm_std * 2)\n```\nMean and std calculated for all train dataset.\n\n##### Model \nSED with tf_efficientnet_b0_ns backbone\n\n##### Augmentations\nSame\n\n##### Training\n* Epochs = 85\n* mixup 0.4, when epochs < 15\n* LR = 0.0008\n* weight_decay = 0.0001\n* dropout = 0.4\n* Loss: Focal BCE Loss\n* Optimizer: Adam (betas=(0.95, 0.999))\n* Scheduler: CosineAnnealingLR\n\n##### Postprocessing\n* Mean-median averaging of predictions\n```\n          full_med = np.median(full_events, axis=1)                    \n          full_mean = np.mean(full_events, axis=1)\n          full_events = np.mean(np.array([full_med, full_mean]), axis=0) \n```\n* Max adder.\n```\n       logits_max = full_events.max(0)\n       for jk in range(full_events.shape[1]):\n           if logits_max[jk] > threshold * 2.5:\n               full_events[:, jk] += threshold * 0.5\n```\n\n## Story of gold medal\nWhen there were 2 days left until the end of the competition, I thought about using a pseudo labeling. And separate my data by 0.8 threshold to remove all noisy records. Only ~15.000 left from ~40.000. Then I trained a new model on this data. Because there were not enough time I trained model on all train data and validate it on some random part and random part of only scored birds. \nAll training and preprocessing parameters was the same with previous.\n\nIt was last day of competition and only 5 submissions. So the first 3 I lost because i forgot to make model.eval() :)\n4th attempt I used with single new model and get **0.79** LB.\nAnd the last one was mean between my best previous attempt and this pseudo labeling new model.\n\nAnd it gave me **0.79** private, when my previous best private was **0.78** ",
      "votes": 32
    },
    {
      "id": 1825327,
      "postDate": "2022-06-19T08:51:06.513Z",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> Thanks for sharing great work. Keep it up.</p>\n<p>Bookmarked and upvoted 💕💕💕</p>",
      "rawMarkdown": "@kolyaforrat Thanks for sharing great work. Keep it up.\n\nBookmarked and upvoted 💕💕💕",
      "votes": 1
    },
    {
      "id": 1802983,
      "postDate": "2022-05-27T10:58:49.313Z",
      "content": "<p>Congrats! on solo gold</p>",
      "rawMarkdown": "Congrats! on solo gold",
      "votes": 1
    },
    {
      "id": 1842005,
      "postDate": "2022-07-03T15:54:53.840Z",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> Congrats! Thanks for sharing great work.</p>",
      "rawMarkdown": "@kolyaforrat Congrats! Thanks for sharing great work."
    },
    {
      "id": 1825270,
      "postDate": "2022-06-19T07:08:40.677Z",
      "content": "<p>great work 😄<br>\njust one quastion <br>\nas i know we use the class weights in our loss function to focus on the classes with small number of samples during training but here you just multiply the predictions with the class weights can you explain it to me theoretically</p>",
      "rawMarkdown": "great work 😄\njust one quastion \nas i know we use the class weights in our loss function to focus on the classes with small number of samples during training but here you just multiply the predictions with the class weights can you explain it to me theoretically",
      "replies": [
        {
          "id": 1825422,
          "postDate": "2022-06-19T10:30:49.783Z",
          "content": "<p>Thanks!</p>\n<p>Weighted loss didn't work for me here. This multiplies it's something like tuning thresholds which other participants does, but I didn't tune thresholds to lower values and just made predictions higher.</p>\n<p>And with this method I didn't tune threshold manually, it always linearly depends of the number of training samples for each class.</p>",
          "rawMarkdown": "Thanks!\n\nWeighted loss didn't work for me here. This multiplies it's something like tuning thresholds which other participants does, but I didn't tune thresholds to lower values and just made predictions higher.\n\nAnd with this method I didn't tune threshold manually, it always linearly depends of the number of training samples for each class.",
          "votes": 1
        },
        {
          "id": 1825516,
          "postDate": "2022-06-19T12:23:38.523Z",
          "content": "<p>Thank you 😄</p>",
          "rawMarkdown": "Thank you 😄"
        },
        {
          "id": 1841456,
          "postDate": "2022-07-03T06:07:51.480Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a>  , sorry for the inconvenience. I tried to do what you mentioned and as shown in the following code, but I did not see any improvement in the validation results. Also, I think that a threshold value should be extracted at the end as well. If it is true, did I rely on the validation set to find it .Thanks again😄</p>\n<pre><code>#params.target_columns contains 152 labels\nweights_dict = {i:int(0) for i in params.target_columns}\nfor targs in train.new_target:\n    targs = targs.split(' ')\n    for targ in targs:\n        if targ != '':\n            weights_dict[targ] += 1\nfor targ in params.target_columns:\n    if targ != '':\n        weights_dict[targ] = 500 / weights_dict[targ]\n        if  weights_dict[targ] &gt; 10:\n             weights_dict[targ] = 10\nweights_list = np.array([weights_dict[x] for x in weights_dict])\n\nparams.mode = \"valid\"\nwith torch.no_grad():\n    all_preds = np.array([])\n    all_targs = np.array([])\n    for n,(xb,yb) in enumerate(tqdm(learn.dls.valid)):\n        preds_dict = learn.model(xb)\n        yb = yb.detach().cpu().numpy()\n        preds = preds_dict['clipwise_output'].detach().cpu().numpy()\n        preds = preds * weights_list  #&lt;---\n        if n == 0:\n            all_preds = preds\n            all_targs = yb\n        else:\n            all_preds = np.vstack((all_preds,preds))\n            all_targs = np.vstack((all_targs,yb))\n\nthreshs = np.linspace(0.0,2,1000)\nbest_score = 0\nfor thresh in tqdm(list(threshs)):\n    score = metrics.f1_score(all_targs,all_preds &gt; thresh, average=\"micro\")  \n    print(\"score\",score)\n    print(\"thresh\",thresh)\n    if score &gt; best_score:\n        best_score = score\nprint(best_score )\n</code></pre>",
          "rawMarkdown": "Hello @kolyaforrat  , sorry for the inconvenience. I tried to do what you mentioned and as shown in the following code, but I did not see any improvement in the validation results. Also, I think that a threshold value should be extracted at the end as well. If it is true, did I rely on the validation set to find it .Thanks again😄\n```\n#params.target_columns contains 152 labels\nweights_dict = {i:int(0) for i in params.target_columns}\nfor targs in train.new_target:\n    targs = targs.split(' ')\n    for targ in targs:\n        if targ != '':\n            weights_dict[targ] += 1\nfor targ in params.target_columns:\n    if targ != '':\n        weights_dict[targ] = 500 / weights_dict[targ]\n        if  weights_dict[targ] > 10:\n             weights_dict[targ] = 10\nweights_list = np.array([weights_dict[x] for x in weights_dict])\n\nparams.mode = \"valid\"\nwith torch.no_grad():\n    all_preds = np.array([])\n    all_targs = np.array([])\n    for n,(xb,yb) in enumerate(tqdm(learn.dls.valid)):\n        preds_dict = learn.model(xb)\n        yb = yb.detach().cpu().numpy()\n        preds = preds_dict['clipwise_output'].detach().cpu().numpy()\n        preds = preds * weights_list  #<---\n        if n == 0:\n            all_preds = preds\n            all_targs = yb\n        else:\n            all_preds = np.vstack((all_preds,preds))\n            all_targs = np.vstack((all_targs,yb))\n\nthreshs = np.linspace(0.0,2,1000)\nbest_score = 0\nfor thresh in tqdm(list(threshs)):\n    score = metrics.f1_score(all_targs,all_preds > thresh, average=\"micro\")  \n    print(\"score\",score)\n    print(\"thresh\",thresh)\n    if score > best_score:\n        best_score = score\nprint(best_score )\n```"
        }
      ]
    },
    {
      "id": 1801762,
      "postDate": "2022-05-26T05:57:55.633Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1825327,
      "author_name": "Muhammad Irfan Azam",
      "author_url": "",
      "post_date": "2022-06-19T08:51:06.513000",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> Thanks for sharing great work. Keep it up.</p>\n<p>Bookmarked and upvoted 💕💕💕</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1802983,
      "author_name": "Neo",
      "author_url": "",
      "post_date": "2022-05-27T10:58:49.313000",
      "content": "<p>Congrats! on solo gold</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1842005,
      "author_name": "Vipin Kumar",
      "author_url": "",
      "post_date": "2022-07-03T15:54:53.840000",
      "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> Congrats! Thanks for sharing great work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1825270,
      "author_name": "RIYADH HAITHAM",
      "author_url": "",
      "post_date": "2022-06-19T07:08:40.677000",
      "content": "<p>great work 😄<br>\njust one quastion <br>\nas i know we use the class weights in our loss function to focus on the classes with small number of samples during training but here you just multiply the predictions with the class weights can you explain it to me theoretically</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1825422,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2022-06-19T10:30:49.783000",
          "content": "<p>Thanks!</p>\n<p>Weighted loss didn't work for me here. This multiplies it's something like tuning thresholds which other participants does, but I didn't tune thresholds to lower values and just made predictions higher.</p>\n<p>And with this method I didn't tune threshold manually, it always linearly depends of the number of training samples for each class.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1825516,
          "author_name": "RIYADH HAITHAM",
          "author_url": "",
          "post_date": "2022-06-19T12:23:38.523000",
          "content": "<p>Thank you 😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1841456,
          "author_name": "RIYADH HAITHAM",
          "author_url": "",
          "post_date": "2022-07-03T06:07:51.480000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a>  , sorry for the inconvenience. I tried to do what you mentioned and as shown in the following code, but I did not see any improvement in the validation results. Also, I think that a threshold value should be extracted at the end as well. If it is true, did I rely on the validation set to find it .Thanks again😄</p>\n<pre><code>#params.target_columns contains 152 labels\nweights_dict = {i:int(0) for i in params.target_columns}\nfor targs in train.new_target:\n    targs = targs.split(' ')\n    for targ in targs:\n        if targ != '':\n            weights_dict[targ] += 1\nfor targ in params.target_columns:\n    if targ != '':\n        weights_dict[targ] = 500 / weights_dict[targ]\n        if  weights_dict[targ] &gt; 10:\n             weights_dict[targ] = 10\nweights_list = np.array([weights_dict[x] for x in weights_dict])\n\nparams.mode = \"valid\"\nwith torch.no_grad():\n    all_preds = np.array([])\n    all_targs = np.array([])\n    for n,(xb,yb) in enumerate(tqdm(learn.dls.valid)):\n        preds_dict = learn.model(xb)\n        yb = yb.detach().cpu().numpy()\n        preds = preds_dict['clipwise_output'].detach().cpu().numpy()\n        preds = preds * weights_list  #&lt;---\n        if n == 0:\n            all_preds = preds\n            all_targs = yb\n        else:\n            all_preds = np.vstack((all_preds,preds))\n            all_targs = np.vstack((all_targs,yb))\n\nthreshs = np.linspace(0.0,2,1000)\nbest_score = 0\nfor thresh in tqdm(list(threshs)):\n    score = metrics.f1_score(all_targs,all_preds &gt; thresh, average=\"micro\")  \n    print(\"score\",score)\n    print(\"thresh\",thresh)\n    if score &gt; best_score:\n        best_score = score\nprint(best_score )\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1801762,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-26T05:57:55.633000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1800868": "It was my first Kaggle competition on which I spent a lot of time. I'm honor to be part of it, big thanks to hosts, Kaggle and all participants.\n\n## First step\n\n1. I started with @kaerunantoka public [notebook](https://www.kaggle.com/code/kaerunantoka/birdclef2022-use-2nd-label-f0) with change mel spectrogram hop length and input shape [224, 512] and add **class weights** to a submission.\nWeights array is 500 divided to amount of each birds species, and clamp it to max value 10. Then I multiply model output to weights array. **0.77** Public LB.\n\n2. I tried this approach with secondary labels with 0.3, 0.4 and 0.5 labels. And with ensembling I'v got **0.79** on LB\n\n##### Augmentations\nFor waveform:\n```\nCompose([OneOf(\n     [AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.011, p=1),\n            NoiseInjection(p=1, max_noise_level=0.04)], p=0.4),                                \n            PitchShift(min_semitones=-4, max_semitones=4, p=0.1),\n            Shift(min_fraction=-0.5, max_fraction=0.5, p=0.1),\n            Gain(min_gain_in_db=-12, max_gain_in_db=12, p=0.2),    \n            Normalize(p=1.) ])\n```\nFor spectrogram:\n ```\ntorchaudio.transforms.FrequencyMasking(24)\ntorchaudio.transforms.TimeMasking(96)\n```\n##### Model \nSED with tf_efficientnet_b0_ns backbone\n\n##### Training\n* Epochs = 30\n* LR = 0.001\n* weight_decay = 0.0001\n* dropout = 0.4\n* Loss: Focal BCE Loss\n* Optimizer: Adam (betas=(0.95, 0.999))\n* Scheduler: CosineAnnealingLR\n\n##### Validation\nI validate models on 7 folds CV with f1-macro metric. Thresholds for f1: [0.5, 0.1, 0.15, 0.20, 0.25, 0.30, 0.35, 0.4, 0.45, 0.5]. Best checkpoints chosen by 0.3 and 0.5 thresholds.\n\n##### What didn't work on this step\n1. Noise reduction\n2. Oversampling rare birds\n3. Trim long silence with noise reduction\n4. Weighted loss\n5. 3 channels input, where 1st channel is power_to_db, 2nd melspectrogram, 3rd normalized melspectrogram\n\n\n## Second step\n1. I manually trim all segments without bird's sounds from the scored audios, except skylar and houfin, on them I processed only 70 records per class. \n2. Splited data to 15 seconds chunks for audios with length less than 1 minute and 30 seconds chunks for more than 1 minute. Got ~40.000 records.\n3. Used weights arrays with 0.25-0.75 power to reduce their impact to inference.\n4. Training on this data with previous pipeline and ensemble with previous models gives me **0.81** on LB\n\n##### Data preprocessing\nSame\n\n##### Augmentations\nSame\n\n##### Model \nSED with tf_efficientnet_b0_ns backbone\n\n##### Training\n* Epochs = 60\n* LR = 0.0013\n* weight_decay = 0.0001\n* dropout = 0.4\n* Loss: Focal BCE Loss\n* Optimizer: Adam (betas=(0.95, 0.999))\n* Scheduler: CosineAnnealingLR\n\n##### What didn't work on this step\n1. PaSST model\n2. AST model\n3. Linear head models\n4. PaSST preprocessing\n\n## Third step\n1. I used similar with AST preprocessing.\n2. Every epoch from all data randomly choose up to 300 records for every class.\n3. Train on random 5 seconds crop, validate on first 5 seconds.\n4. Secondary labels 0.4\n5. Class weights array clamped with max value 8. And used with power 0.6. \n6. Used mixup 0.4 for first 15 epochs, mixup 0.07 for 16-22 epochs, and no mixup for others.\n7. Ensembling of this approach with previous gives me **0.82** LB\n\n##### Data preprocessing\n```\nwaveform, sr = ta.load(filename)\nwaveform = crop_or_pad(waveform, sr=SR, mode=self.mode)\nwaveform = waveform - waveform.mean()\nwaveform = torch.tensor(self.wave_transforms(samples=waveform[0].numpy(), sample_rate=SR)).unsqueeze(0)\nfbank = ta.compliance.kaldi.fbank(waveform, htk_compat=True,          sample_frequency=SR, use_energy=False, window_type='hanning',                 num_mel_bins=self.melbins, dither=0.0, frame_shift=9.7)\nfbank = (fbank - self.norm_mean) / (self.norm_std * 2)\n```\nMean and std calculated for all train dataset.\n\n##### Model \nSED with tf_efficientnet_b0_ns backbone\n\n##### Augmentations\nSame\n\n##### Training\n* Epochs = 85\n* mixup 0.4, when epochs < 15\n* LR = 0.0008\n* weight_decay = 0.0001\n* dropout = 0.4\n* Loss: Focal BCE Loss\n* Optimizer: Adam (betas=(0.95, 0.999))\n* Scheduler: CosineAnnealingLR\n\n##### Postprocessing\n* Mean-median averaging of predictions\n```\n          full_med = np.median(full_events, axis=1)                    \n          full_mean = np.mean(full_events, axis=1)\n          full_events = np.mean(np.array([full_med, full_mean]), axis=0) \n```\n* Max adder.\n```\n       logits_max = full_events.max(0)\n       for jk in range(full_events.shape[1]):\n           if logits_max[jk] > threshold * 2.5:\n               full_events[:, jk] += threshold * 0.5\n```\n\n## Story of gold medal\nWhen there were 2 days left until the end of the competition, I thought about using a pseudo labeling. And separate my data by 0.8 threshold to remove all noisy records. Only ~15.000 left from ~40.000. Then I trained a new model on this data. Because there were not enough time I trained model on all train data and validate it on some random part and random part of only scored birds. \nAll training and preprocessing parameters was the same with previous.\n\nIt was last day of competition and only 5 submissions. So the first 3 I lost because i forgot to make model.eval() :)\n4th attempt I used with single new model and get **0.79** LB.\nAnd the last one was mean between my best previous attempt and this pseudo labeling new model.\n\nAnd it gave me **0.79** private, when my previous best private was **0.78** ",
    "1825327": "@kolyaforrat Thanks for sharing great work. Keep it up.\n\nBookmarked and upvoted 💕💕💕",
    "1802983": "Congrats! on solo gold",
    "1842005": "@kolyaforrat Congrats! Thanks for sharing great work.",
    "1825270": "great work 😄\njust one quastion \nas i know we use the class weights in our loss function to focus on the classes with small number of samples during training but here you just multiply the predictions with the class weights can you explain it to me theoretically",
    "1801762": ""
  }
}