{
  "id": 328649,
  "title": "35th Experiment",
  "url": "/competitions/birdclef-2022/writeups/sqrt4kaido-35th-experiment",
  "author_name": "",
  "post_date": "2022-06-02T10:09:00.363Z",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi. I am disappointed that I could not keep the top position until the end of the competition.<br>\nI would like to leave some notes about the experiments I did during the competition.<br>\n(日本語版は<a href=\"https://sqrt4kaido.hatenablog.com/entry/2022/05/29/034449?_ga=2.139101755.906933634.1654160869-182752224.1648474566\" target=\"_blank\">こちら</a>)</p>\n<h2>Feature</h2>\n<p>I considered methods more robust to noise than the melspectrogram, or other conversion methods from the melspectrogram.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-per-channel-energy-normalization\" target=\"_blank\">pcen</a></li>\n</ul>\n<p>I experimented with the image with low noise and high expectations, but the score did not improve.</p>\n<ul>\n<li>linear spectrogram</li>\n</ul>\n<p>Thinking that birdcall features might exist in frequency bands that are difficult for the human ear to hear, I created an image using an unbiased moving average (which I call a linear spectrogram). However, the score did not improve either.</p>\n<p><img src=\"https://raw.githubusercontent.com/root4kaido/kaggle-discussion-image/main/BIRDCLEF2022/mel.png\" alt=\"image\"></p>\n<p>This is just a tip, but the melspectrogram with torchaudio looks slightly different from the mel spectrogram with librosa.<br>\nThe same thing happens when taking logs. Please wait to upload the consistent code on kaggle notebook.</p>\n<p>In the end, the best scores were obtained using the regular melspectrogram. pcen and linear spectrograms were tried in combination with the melspectrogram and by shifting the frequency band used for each channel, but to no avail.</p>\n<h2>Model</h2>\n<p>I have considered/implemented three models.</p>\n<ul>\n<li>PANNs-based model</li>\n</ul>\n<p>This is a <a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">familiar model</a> for bird competitions. The Linear layer after passing through the backbone was changed to LSTM in order to capture more time-series features. This led to a slight improvement in score.</p>\n<ul>\n<li><a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">STFT Transformer</a></li>\n</ul>\n<p>CPMP invented this in last year's bird competition.<br>\nHe scored highly using this method in that competition, but in my experiments, I could not surpass PANNs' score.</p>\n<ul>\n<li>STFT Swin SED model</li>\n</ul>\n<p>This is a model that I wanted to implement from the beginning of the competition. It is based on the above STFT Transformer, but is capable of outputting class classification results for each frame.<br>\nSince the above STFT Transformer cannot output frame-by-frame classification results, I thought that adding a mechanism more suitable for SED tasks with a similar architecture to the STFT Transformer would lead to higher scores.<br>\nI changed the method of embedding patches to the STFT Transformer, based on <a href=\"https://arxiv.org/abs/2202.00874\" target=\"_blank\">the paper</a> on the SED model using the Swin Transformer.<br>\nHowever, I could not exceed the PANNs score.</p>\n<h2>Class imbalance</h2>\n<ul>\n<li>over sampling</li>\n</ul>\n<p>There are two methods that have improved scores.</p>\n<p>Oversampling was applied to wavs of species with fewer than 20 wavs. To prevent overfitting of the oversampled wavs, augmentation was strongly applied only to the oversampled wavs. This improved the score slightly.</p>\n<p>I used A method called Context-rich Minority Oversampling (<a href=\"https://arxiv.org/abs/2112.00412\" target=\"_blank\">CMO</a>). This method oversamples a small number of classes at the time of CutMix, and was effective when applied to the mel-spectrogram. I tried different cutting methods, such as CutMixing only in the time direction, but the simple CutMixing was the best.</p>\n<ul>\n<li>loss</li>\n</ul>\n<p>In addition to focal loss, I also examined several effective losses for class imbalance. I implemented <a href=\"https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf\" target=\"_blank\">CBLoss</a>, <a href=\"https://arxiv.org/abs/2009.14119\" target=\"_blank\">ASL</a>, and <a href=\"https://openaccess.thecvf.com/content/ICCV2021/html/Park_Influence-Balanced_Loss_for_Imbalanced_Visual_Classification_ICCV_2021_paper.html\" target=\"_blank\">IBLoss</a>, but in the end I could not achieve an accuracy higher than that of focal loss. Among them, the method called IBloss, which reduces the weight of samples near the decision boundary from half of the total epoch, was implemented with high expectations. However, although the CV score increased, the Public score did not improve.</p>\n<h2>Local metric</h2>\n<p>I used the following code to calculate CV scores, but in the end I could not correlate them with LB at all. Therefore, my strategy was to calculate this metric score for the 21 classes to be evaluated and the f1 scores (samples) for all 152 classes at various thresholds, and submit the results if they were above the majority.</p>\n<pre><code>def calc_tpr(true, pred, threshold):\n    true = (true &gt; 0.5) * 1\n    pred = (pred &gt; threshold) * 1\n    tp = np.sum(true * pred)\n    fn = np.sum(true * (1-pred))\n    tpr = tp / (tp + fn + 1e-6)\n    return tpr\n\ndef calc_tnr(true, pred, threshold):\n    true = (true &gt; 0.5) * 1\n    pred = (pred &gt; threshold) * 1\n    tn = np.sum((1-true) * (1-pred))\n    fp = np.sum((1-true) * pred)\n    tnr = tn / (tn + fp + 1e-6)\n    return tnr\n\nclass Metric():\n    def __init__(self):\n        self.tgt_index = list(map(lambda x: CFG.target_columns.index(x), CFG.test_target_columns))\n\n    def calc_tgt_f1_score(self, gt_arr, pred_arr, th, epsilon=1e-9):\n        scores = []\n        for i in self.tgt_index:\n            a = list(gt_arr[:, i])\n            uni_, counts = np.unique(a, return_counts=True)\n        #     posi_weight = counts[0] / len(a)\n        #     nega_weight = counts[1] / len(a)\n            posi_weight = 0.5\n            nega_weight = 0.5\n\n            tpr = calc_tpr(gt_arr[:, i], pred_arr[:, i], th)\n            tnr = calc_tnr(gt_arr[:, i], pred_arr[:, i], th)\n\n            score = posi_weight * tpr + nega_weight * tnr\n            scores.append(score)\n\n        return np.mean(scores)\n</code></pre>\n<h2>Inference</h2>\n<p>A much lower threshold seemed to be appropriate for this competition, I thought there are two reasons for this.</p>\n<ul>\n<li>It is important to increase the TPR even if the FPR is increased since the number of positive classes is quite small in the test data.</li>\n<li>The decision boundaries are quite narrow, especially for the few classes, and the threshold must be lowered considerably in order to be judged as positive.</li>\n</ul>\n<p>Because of the second reason, I thought the majority class would be able to detect a positive class even with a higher threshold value. Therefore, I determined the threshold based on the number of data per class. For example, the threshold for the class with the most training data was set at 0.3, the threshold for the class with the least training data was set at 0.05, and the rest were scaled linearly based on the number of data.</p>\n<p>There is one more innovation. Initially, when a bird was found in a wav, post-processing was performed by adding 0.2 to the prediction probability of that class for the entire wav. However,  for example, the threshold is set to 0.05, adding 0.2 would mean that a bird of the target class is detected in a total of that wav. Therefore, in my reasoning, I only made predictions on a per-sound-file basis for the 19 bird species for which there was little training data (i.e., the threshold value did not exceed 0.2) of the 21 species to be evaluated.</p>\n<p>Various other post-processing methods were also tried, but they did not improve the score. The following are some examples.</p>\n<ul>\n<li>Detecting the presence of something like a monophonic sound in the mel-spectrogram using image processing. If no birds are detected, it is assumed that no birds are singing.</li>\n<li>Use co-occurrence matrices to increase the probability of predicting birds that may be singing with the bird with the highest probability of prediction.</li>\n</ul>\n<h2>Other</h2>\n<p>I created a model targeting latitude and longitude, and for species found only in Hawaii, I determined that the target bird was not singing when the sound was determined to be from outside of Hawaii.<br>\nThe model by itself produced a public score of about 0.6, but combining it with an existing model did not improve the accuracy.</p>\n<p>For pseudo labels, I tried adding additional secondary_labels using 5fold oof, and conversely, tried removing suspicious ones.<br>\nNone of them worked. I have seen some reports of improved accuracy with manual labeling, so perhaps the original model was not accurate enough.</p>\n<p>That's all. Thank you very much for reading.</p>",
  "messages": [
    {
      "id": "1808999",
      "postDate": "06/02/2022 10:04:27",
      "content": "<p>Hi. I am disappointed that I could not keep the top position until the end of the competition.<br>\nI would like to leave some notes about the experiments I did during the competition.<br>\n(日本語版は<a href=\"https://sqrt4kaido.hatenablog.com/entry/2022/05/29/034449?_ga=2.139101755.906933634.1654160869-182752224.1648474566\" target=\"_blank\">こちら</a>)</p>\n<h2>Feature</h2>\n<p>I considered methods more robust to noise than the melspectrogram, or other conversion methods from the melspectrogram.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-per-channel-energy-normalization\" target=\"_blank\">pcen</a></li>\n</ul>\n<p>I experimented with the image with low noise and high expectations, but the score did not improve.</p>\n<ul>\n<li>linear spectrogram</li>\n</ul>\n<p>Thinking that birdcall features might exist in frequency bands that are difficult for the human ear to hear, I created an image using an unbiased moving average (which I call a linear spectrogram). However, the score did not improve either.</p>\n<p><img src=\"https://raw.githubusercontent.com/root4kaido/kaggle-discussion-image/main/BIRDCLEF2022/mel.png\" alt=\"image\"></p>\n<p>This is just a tip, but the melspectrogram with torchaudio looks slightly different from the mel spectrogram with librosa.<br>\nThe same thing happens when taking logs. Please wait to upload the consistent code on kaggle notebook.</p>\n<p>In the end, the best scores were obtained using the regular melspectrogram. pcen and linear spectrograms were tried in combination with the melspectrogram and by shifting the frequency band used for each channel, but to no avail.</p>\n<h2>Model</h2>\n<p>I have considered/implemented three models.</p>\n<ul>\n<li>PANNs-based model</li>\n</ul>\n<p>This is a <a href=\"https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">familiar model</a> for bird competitions. The Linear layer after passing through the backbone was changed to LSTM in order to capture more time-series features. This led to a slight improvement in score.</p>\n<ul>\n<li><a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">STFT Transformer</a></li>\n</ul>\n<p>CPMP invented this in last year's bird competition.<br>\nHe scored highly using this method in that competition, but in my experiments, I could not surpass PANNs' score.</p>\n<ul>\n<li>STFT Swin SED model</li>\n</ul>\n<p>This is a model that I wanted to implement from the beginning of the competition. It is based on the above STFT Transformer, but is capable of outputting class classification results for each frame.<br>\nSince the above STFT Transformer cannot output frame-by-frame classification results, I thought that adding a mechanism more suitable for SED tasks with a similar architecture to the STFT Transformer would lead to higher scores.<br>\nI changed the method of embedding patches to the STFT Transformer, based on <a href=\"https://arxiv.org/abs/2202.00874\" target=\"_blank\">the paper</a> on the SED model using the Swin Transformer.<br>\nHowever, I could not exceed the PANNs score.</p>\n<h2>Class imbalance</h2>\n<ul>\n<li>over sampling</li>\n</ul>\n<p>There are two methods that have improved scores.</p>\n<p>Oversampling was applied to wavs of species with fewer than 20 wavs. To prevent overfitting of the oversampled wavs, augmentation was strongly applied only to the oversampled wavs. This improved the score slightly.</p>\n<p>I used A method called Context-rich Minority Oversampling (<a href=\"https://arxiv.org/abs/2112.00412\" target=\"_blank\">CMO</a>). This method oversamples a small number of classes at the time of CutMix, and was effective when applied to the mel-spectrogram. I tried different cutting methods, such as CutMixing only in the time direction, but the simple CutMixing was the best.</p>\n<ul>\n<li>loss</li>\n</ul>\n<p>In addition to focal loss, I also examined several effective losses for class imbalance. I implemented <a href=\"https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf\" target=\"_blank\">CBLoss</a>, <a href=\"https://arxiv.org/abs/2009.14119\" target=\"_blank\">ASL</a>, and <a href=\"https://openaccess.thecvf.com/content/ICCV2021/html/Park_Influence-Balanced_Loss_for_Imbalanced_Visual_Classification_ICCV_2021_paper.html\" target=\"_blank\">IBLoss</a>, but in the end I could not achieve an accuracy higher than that of focal loss. Among them, the method called IBloss, which reduces the weight of samples near the decision boundary from half of the total epoch, was implemented with high expectations. However, although the CV score increased, the Public score did not improve.</p>\n<h2>Local metric</h2>\n<p>I used the following code to calculate CV scores, but in the end I could not correlate them with LB at all. Therefore, my strategy was to calculate this metric score for the 21 classes to be evaluated and the f1 scores (samples) for all 152 classes at various thresholds, and submit the results if they were above the majority.</p>\n<pre><code>def calc_tpr(true, pred, threshold):\n    true = (true &gt; 0.5) * 1\n    pred = (pred &gt; threshold) * 1\n    tp = np.sum(true * pred)\n    fn = np.sum(true * (1-pred))\n    tpr = tp / (tp + fn + 1e-6)\n    return tpr\n\ndef calc_tnr(true, pred, threshold):\n    true = (true &gt; 0.5) * 1\n    pred = (pred &gt; threshold) * 1\n    tn = np.sum((1-true) * (1-pred))\n    fp = np.sum((1-true) * pred)\n    tnr = tn / (tn + fp + 1e-6)\n    return tnr\n\nclass Metric():\n    def __init__(self):\n        self.tgt_index = list(map(lambda x: CFG.target_columns.index(x), CFG.test_target_columns))\n\n    def calc_tgt_f1_score(self, gt_arr, pred_arr, th, epsilon=1e-9):\n        scores = []\n        for i in self.tgt_index:\n            a = list(gt_arr[:, i])\n            uni_, counts = np.unique(a, return_counts=True)\n        #     posi_weight = counts[0] / len(a)\n        #     nega_weight = counts[1] / len(a)\n            posi_weight = 0.5\n            nega_weight = 0.5\n\n            tpr = calc_tpr(gt_arr[:, i], pred_arr[:, i], th)\n            tnr = calc_tnr(gt_arr[:, i], pred_arr[:, i], th)\n\n            score = posi_weight * tpr + nega_weight * tnr\n            scores.append(score)\n\n        return np.mean(scores)\n</code></pre>\n<h2>Inference</h2>\n<p>A much lower threshold seemed to be appropriate for this competition, I thought there are two reasons for this.</p>\n<ul>\n<li>It is important to increase the TPR even if the FPR is increased since the number of positive classes is quite small in the test data.</li>\n<li>The decision boundaries are quite narrow, especially for the few classes, and the threshold must be lowered considerably in order to be judged as positive.</li>\n</ul>\n<p>Because of the second reason, I thought the majority class would be able to detect a positive class even with a higher threshold value. Therefore, I determined the threshold based on the number of data per class. For example, the threshold for the class with the most training data was set at 0.3, the threshold for the class with the least training data was set at 0.05, and the rest were scaled linearly based on the number of data.</p>\n<p>There is one more innovation. Initially, when a bird was found in a wav, post-processing was performed by adding 0.2 to the prediction probability of that class for the entire wav. However,  for example, the threshold is set to 0.05, adding 0.2 would mean that a bird of the target class is detected in a total of that wav. Therefore, in my reasoning, I only made predictions on a per-sound-file basis for the 19 bird species for which there was little training data (i.e., the threshold value did not exceed 0.2) of the 21 species to be evaluated.</p>\n<p>Various other post-processing methods were also tried, but they did not improve the score. The following are some examples.</p>\n<ul>\n<li>Detecting the presence of something like a monophonic sound in the mel-spectrogram using image processing. If no birds are detected, it is assumed that no birds are singing.</li>\n<li>Use co-occurrence matrices to increase the probability of predicting birds that may be singing with the bird with the highest probability of prediction.</li>\n</ul>\n<h2>Other</h2>\n<p>I created a model targeting latitude and longitude, and for species found only in Hawaii, I determined that the target bird was not singing when the sound was determined to be from outside of Hawaii.<br>\nThe model by itself produced a public score of about 0.6, but combining it with an existing model did not improve the accuracy.</p>\n<p>For pseudo labels, I tried adding additional secondary_labels using 5fold oof, and conversely, tried removing suspicious ones.<br>\nNone of them worked. I have seen some reports of improved accuracy with manual labeling, so perhaps the original model was not accurate enough.</p>\n<p>That's all. Thank you very much for reading.</p>",
      "rawMarkdown": "Hi. I am disappointed that I could not keep the top position until the end of the competition.\nI would like to leave some notes about the experiments I did during the competition.\n(日本語版は[こちら](https://sqrt4kaido.hatenablog.com/entry/2022/05/29/034449?_ga=2.139101755.906933634.1654160869-182752224.1648474566))\n\n## Feature\nI considered methods more robust to noise than the melspectrogram, or other conversion methods from the melspectrogram.\n\n- [pcen](https://www.kaggle.com/code/tatamikenn/birdclef22-per-channel-energy-normalization)\n\nI experimented with the image with low noise and high expectations, but the score did not improve.\n\n- linear spectrogram\n\nThinking that birdcall features might exist in frequency bands that are difficult for the human ear to hear, I created an image using an unbiased moving average (which I call a linear spectrogram). However, the score did not improve either.\n\n![image](https://raw.githubusercontent.com/root4kaido/kaggle-discussion-image/main/BIRDCLEF2022/mel.png)\n\nThis is just a tip, but the melspectrogram with torchaudio looks slightly different from the mel spectrogram with librosa.\nThe same thing happens when taking logs. Please wait to upload the consistent code on kaggle notebook.\n\nIn the end, the best scores were obtained using the regular melspectrogram. pcen and linear spectrograms were tried in combination with the melspectrogram and by shifting the frequency band used for each channel, but to no avail.\n\n## Model\n\nI have considered/implemented three models.\n\n- PANNs-based model\n\nThis is a [familiar model](https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection) for bird competitions. The Linear layer after passing through the backbone was changed to LSTM in order to capture more time-series features. This led to a slight improvement in score.\n\n- [STFT Transformer](https://github.com/jfpuget/STFT_Transformer)\n\nCPMP invented this in last year's bird competition.\nHe scored highly using this method in that competition, but in my experiments, I could not surpass PANNs' score.\n\n- STFT Swin SED model\n\nThis is a model that I wanted to implement from the beginning of the competition. It is based on the above STFT Transformer, but is capable of outputting class classification results for each frame.\nSince the above STFT Transformer cannot output frame-by-frame classification results, I thought that adding a mechanism more suitable for SED tasks with a similar architecture to the STFT Transformer would lead to higher scores.\nI changed the method of embedding patches to the STFT Transformer, based on [the paper](https://arxiv.org/abs/2202.00874) on the SED model using the Swin Transformer.\nHowever, I could not exceed the PANNs score.\n\n## Class imbalance\n\n- over sampling\n\nThere are two methods that have improved scores.\n\nOversampling was applied to wavs of species with fewer than 20 wavs. To prevent overfitting of the oversampled wavs, augmentation was strongly applied only to the oversampled wavs. This improved the score slightly.\n\nI used A method called Context-rich Minority Oversampling ([CMO](https://arxiv.org/abs/2112.00412)). This method oversamples a small number of classes at the time of CutMix, and was effective when applied to the mel-spectrogram. I tried different cutting methods, such as CutMixing only in the time direction, but the simple CutMixing was the best.\n\n- loss\n\nIn addition to focal loss, I also examined several effective losses for class imbalance. I implemented [CBLoss](https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf), [ASL](https://arxiv.org/abs/2009.14119), and [IBLoss](https://openaccess.thecvf.com/content/ICCV2021/html/Park_Influence-Balanced_Loss_for_Imbalanced_Visual_Classification_ICCV_2021_paper.html), but in the end I could not achieve an accuracy higher than that of focal loss. Among them, the method called IBloss, which reduces the weight of samples near the decision boundary from half of the total epoch, was implemented with high expectations. However, although the CV score increased, the Public score did not improve.\n\n## Local metric\n\nI used the following code to calculate CV scores, but in the end I could not correlate them with LB at all. Therefore, my strategy was to calculate this metric score for the 21 classes to be evaluated and the f1 scores (samples) for all 152 classes at various thresholds, and submit the results if they were above the majority.\n\n```\ndef calc_tpr(true, pred, threshold):\n    true = (true > 0.5) * 1\n    pred = (pred > threshold) * 1\n    tp = np.sum(true * pred)\n    fn = np.sum(true * (1-pred))\n    tpr = tp / (tp + fn + 1e-6)\n    return tpr\n    \ndef calc_tnr(true, pred, threshold):\n    true = (true > 0.5) * 1\n    pred = (pred > threshold) * 1\n    tn = np.sum((1-true) * (1-pred))\n    fp = np.sum((1-true) * pred)\n    tnr = tn / (tn + fp + 1e-6)\n    return tnr\n\nclass Metric():\n    def __init__(self):\n        self.tgt_index = list(map(lambda x: CFG.target_columns.index(x), CFG.test_target_columns))\n        \n    def calc_tgt_f1_score(self, gt_arr, pred_arr, th, epsilon=1e-9):\n        scores = []\n        for i in self.tgt_index:\n            a = list(gt_arr[:, i])\n            uni_, counts = np.unique(a, return_counts=True)\n        #     posi_weight = counts[0] / len(a)\n        #     nega_weight = counts[1] / len(a)\n            posi_weight = 0.5\n            nega_weight = 0.5\n\n            tpr = calc_tpr(gt_arr[:, i], pred_arr[:, i], th)\n            tnr = calc_tnr(gt_arr[:, i], pred_arr[:, i], th)\n\n            score = posi_weight * tpr + nega_weight * tnr\n            scores.append(score)\n\n        return np.mean(scores)\n```\n\n## Inference\n\nA much lower threshold seemed to be appropriate for this competition, I thought there are two reasons for this.\n\n- It is important to increase the TPR even if the FPR is increased since the number of positive classes is quite small in the test data.\n- The decision boundaries are quite narrow, especially for the few classes, and the threshold must be lowered considerably in order to be judged as positive.\n\nBecause of the second reason, I thought the majority class would be able to detect a positive class even with a higher threshold value. Therefore, I determined the threshold based on the number of data per class. For example, the threshold for the class with the most training data was set at 0.3, the threshold for the class with the least training data was set at 0.05, and the rest were scaled linearly based on the number of data.\n\nThere is one more innovation. Initially, when a bird was found in a wav, post-processing was performed by adding 0.2 to the prediction probability of that class for the entire wav. However,  for example, the threshold is set to 0.05, adding 0.2 would mean that a bird of the target class is detected in a total of that wav. Therefore, in my reasoning, I only made predictions on a per-sound-file basis for the 19 bird species for which there was little training data (i.e., the threshold value did not exceed 0.2) of the 21 species to be evaluated.\n\nVarious other post-processing methods were also tried, but they did not improve the score. The following are some examples.\n\n- Detecting the presence of something like a monophonic sound in the mel-spectrogram using image processing. If no birds are detected, it is assumed that no birds are singing.\n- Use co-occurrence matrices to increase the probability of predicting birds that may be singing with the bird with the highest probability of prediction.\n\n## Other\n\nI created a model targeting latitude and longitude, and for species found only in Hawaii, I determined that the target bird was not singing when the sound was determined to be from outside of Hawaii.\nThe model by itself produced a public score of about 0.6, but combining it with an existing model did not improve the accuracy.\n\nFor pseudo labels, I tried adding additional secondary_labels using 5fold oof, and conversely, tried removing suspicious ones.\nNone of them worked. I have seen some reports of improved accuracy with manual labeling, so perhaps the original model was not accurate enough.\n\n\nThat's all. Thank you very much for reading.",
      "votes": null
    },
    {
      "id": "1809659",
      "postDate": "06/02/2022 23:26:01",
      "content": "<p>Not being on top doesn't mean failure. Those that see Only the results/success can't see the whole process. </p>\n<p>You did great, Thanks for sharing your worthy experiment.</p>",
      "rawMarkdown": "Not being on top doesn't mean failure. Those that see Only the results/success can't see the whole process. \n\nYou did great, Thanks for sharing your worthy experiment.",
      "votes": null
    },
    {
      "id": "1810159",
      "postDate": "06/03/2022 10:06:55",
      "content": "<p>A notebook matching librosa and torchaudio's melspectrogram is now available.</p>\n<p><a href=\"https://www.kaggle.com/code/nomorevotch/create-the-same-mel-from-librosa-and-torchaudio\" target=\"_blank\">https://www.kaggle.com/code/nomorevotch/create-the-same-mel-from-librosa-and-torchaudio</a></p>",
      "rawMarkdown": "A notebook matching librosa and torchaudio's melspectrogram is now available.\n\nhttps://www.kaggle.com/code/nomorevotch/create-the-same-mel-from-librosa-and-torchaudio",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1809659,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "06/02/2022 23:26:01",
      "content": "<p>Not being on top doesn't mean failure. Those that see Only the results/success can't see the whole process. </p>\n<p>You did great, Thanks for sharing your worthy experiment.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1810159,
      "author_name": "nomorevotch",
      "author_url": "",
      "post_date": "06/03/2022 10:06:55",
      "content": "<p>A notebook matching librosa and torchaudio's melspectrogram is now available.</p>\n<p><a href=\"https://www.kaggle.com/code/nomorevotch/create-the-same-mel-from-librosa-and-torchaudio\" target=\"_blank\">https://www.kaggle.com/code/nomorevotch/create-the-same-mel-from-librosa-and-torchaudio</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1808999": "Hi. I am disappointed that I could not keep the top position until the end of the competition.\nI would like to leave some notes about the experiments I did during the competition.\n(日本語版は[こちら](https://sqrt4kaido.hatenablog.com/entry/2022/05/29/034449?_ga=2.139101755.906933634.1654160869-182752224.1648474566))\n\n## Feature\nI considered methods more robust to noise than the melspectrogram, or other conversion methods from the melspectrogram.\n\n- [pcen](https://www.kaggle.com/code/tatamikenn/birdclef22-per-channel-energy-normalization)\n\nI experimented with the image with low noise and high expectations, but the score did not improve.\n\n- linear spectrogram\n\nThinking that birdcall features might exist in frequency bands that are difficult for the human ear to hear, I created an image using an unbiased moving average (which I call a linear spectrogram). However, the score did not improve either.\n\n![image](https://raw.githubusercontent.com/root4kaido/kaggle-discussion-image/main/BIRDCLEF2022/mel.png)\n\nThis is just a tip, but the melspectrogram with torchaudio looks slightly different from the mel spectrogram with librosa.\nThe same thing happens when taking logs. Please wait to upload the consistent code on kaggle notebook.\n\nIn the end, the best scores were obtained using the regular melspectrogram. pcen and linear spectrograms were tried in combination with the melspectrogram and by shifting the frequency band used for each channel, but to no avail.\n\n## Model\n\nI have considered/implemented three models.\n\n- PANNs-based model\n\nThis is a [familiar model](https://www.kaggle.com/code/hidehisaarai1213/introduction-to-sound-event-detection) for bird competitions. The Linear layer after passing through the backbone was changed to LSTM in order to capture more time-series features. This led to a slight improvement in score.\n\n- [STFT Transformer](https://github.com/jfpuget/STFT_Transformer)\n\nCPMP invented this in last year's bird competition.\nHe scored highly using this method in that competition, but in my experiments, I could not surpass PANNs' score.\n\n- STFT Swin SED model\n\nThis is a model that I wanted to implement from the beginning of the competition. It is based on the above STFT Transformer, but is capable of outputting class classification results for each frame.\nSince the above STFT Transformer cannot output frame-by-frame classification results, I thought that adding a mechanism more suitable for SED tasks with a similar architecture to the STFT Transformer would lead to higher scores.\nI changed the method of embedding patches to the STFT Transformer, based on [the paper](https://arxiv.org/abs/2202.00874) on the SED model using the Swin Transformer.\nHowever, I could not exceed the PANNs score.\n\n## Class imbalance\n\n- over sampling\n\nThere are two methods that have improved scores.\n\nOversampling was applied to wavs of species with fewer than 20 wavs. To prevent overfitting of the oversampled wavs, augmentation was strongly applied only to the oversampled wavs. This improved the score slightly.\n\nI used A method called Context-rich Minority Oversampling ([CMO](https://arxiv.org/abs/2112.00412)). This method oversamples a small number of classes at the time of CutMix, and was effective when applied to the mel-spectrogram. I tried different cutting methods, such as CutMixing only in the time direction, but the simple CutMixing was the best.\n\n- loss\n\nIn addition to focal loss, I also examined several effective losses for class imbalance. I implemented [CBLoss](https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf), [ASL](https://arxiv.org/abs/2009.14119), and [IBLoss](https://openaccess.thecvf.com/content/ICCV2021/html/Park_Influence-Balanced_Loss_for_Imbalanced_Visual_Classification_ICCV_2021_paper.html), but in the end I could not achieve an accuracy higher than that of focal loss. Among them, the method called IBloss, which reduces the weight of samples near the decision boundary from half of the total epoch, was implemented with high expectations. However, although the CV score increased, the Public score did not improve.\n\n## Local metric\n\nI used the following code to calculate CV scores, but in the end I could not correlate them with LB at all. Therefore, my strategy was to calculate this metric score for the 21 classes to be evaluated and the f1 scores (samples) for all 152 classes at various thresholds, and submit the results if they were above the majority.\n\n```\ndef calc_tpr(true, pred, threshold):\n    true = (true > 0.5) * 1\n    pred = (pred > threshold) * 1\n    tp = np.sum(true * pred)\n    fn = np.sum(true * (1-pred))\n    tpr = tp / (tp + fn + 1e-6)\n    return tpr\n    \ndef calc_tnr(true, pred, threshold):\n    true = (true > 0.5) * 1\n    pred = (pred > threshold) * 1\n    tn = np.sum((1-true) * (1-pred))\n    fp = np.sum((1-true) * pred)\n    tnr = tn / (tn + fp + 1e-6)\n    return tnr\n\nclass Metric():\n    def __init__(self):\n        self.tgt_index = list(map(lambda x: CFG.target_columns.index(x), CFG.test_target_columns))\n        \n    def calc_tgt_f1_score(self, gt_arr, pred_arr, th, epsilon=1e-9):\n        scores = []\n        for i in self.tgt_index:\n            a = list(gt_arr[:, i])\n            uni_, counts = np.unique(a, return_counts=True)\n        #     posi_weight = counts[0] / len(a)\n        #     nega_weight = counts[1] / len(a)\n            posi_weight = 0.5\n            nega_weight = 0.5\n\n            tpr = calc_tpr(gt_arr[:, i], pred_arr[:, i], th)\n            tnr = calc_tnr(gt_arr[:, i], pred_arr[:, i], th)\n\n            score = posi_weight * tpr + nega_weight * tnr\n            scores.append(score)\n\n        return np.mean(scores)\n```\n\n## Inference\n\nA much lower threshold seemed to be appropriate for this competition, I thought there are two reasons for this.\n\n- It is important to increase the TPR even if the FPR is increased since the number of positive classes is quite small in the test data.\n- The decision boundaries are quite narrow, especially for the few classes, and the threshold must be lowered considerably in order to be judged as positive.\n\nBecause of the second reason, I thought the majority class would be able to detect a positive class even with a higher threshold value. Therefore, I determined the threshold based on the number of data per class. For example, the threshold for the class with the most training data was set at 0.3, the threshold for the class with the least training data was set at 0.05, and the rest were scaled linearly based on the number of data.\n\nThere is one more innovation. Initially, when a bird was found in a wav, post-processing was performed by adding 0.2 to the prediction probability of that class for the entire wav. However,  for example, the threshold is set to 0.05, adding 0.2 would mean that a bird of the target class is detected in a total of that wav. Therefore, in my reasoning, I only made predictions on a per-sound-file basis for the 19 bird species for which there was little training data (i.e., the threshold value did not exceed 0.2) of the 21 species to be evaluated.\n\nVarious other post-processing methods were also tried, but they did not improve the score. The following are some examples.\n\n- Detecting the presence of something like a monophonic sound in the mel-spectrogram using image processing. If no birds are detected, it is assumed that no birds are singing.\n- Use co-occurrence matrices to increase the probability of predicting birds that may be singing with the bird with the highest probability of prediction.\n\n## Other\n\nI created a model targeting latitude and longitude, and for species found only in Hawaii, I determined that the target bird was not singing when the sound was determined to be from outside of Hawaii.\nThe model by itself produced a public score of about 0.6, but combining it with an existing model did not improve the accuracy.\n\nFor pseudo labels, I tried adding additional secondary_labels using 5fold oof, and conversely, tried removing suspicious ones.\nNone of them worked. I have seen some reports of improved accuracy with manual labeling, so perhaps the original model was not accurate enough.\n\n\nThat's all. Thank you very much for reading.",
    "1809659": "Not being on top doesn't mean failure. Those that see Only the results/success can't see the whole process. \n\nYou did great, Thanks for sharing your worthy experiment.",
    "1810159": "A notebook matching librosa and torchaudio's melspectrogram is now available.\n\nhttps://www.kaggle.com/code/nomorevotch/create-the-same-mel-from-librosa-and-torchaudio"
  },
  "source": "meta"
}