{
  "id": 326990,
  "title": "Few words about 14th place solution",
  "url": "/competitions/birdclef-2022/discussion/326990",
  "author_name": "anthony",
  "post_date": "2022-05-25T07:23:16.765000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to everybody who organized this competition. Congratulations to all who received medals.<br>\nSeems that results of this competition aren't make a lot of sense because the host solution is the best.</p>\n<p>Few words about my trying, hope you find it interesting or even useful.</p>\n<p><strong>Metrics</strong><br>\nI start with macro Fb-score which seems are suitable for case of noisy labels. I chose quite big beta - 16 to boost recall.<br>\nAfter reading some comments of hosts about the metric I switched to:<br>\n<code>M = (1/N) * (1/21) * ∑∑( (TP/(TP+FN) + TN/(TN+FP))/2 )</code></p>\n<p>For any amount of training classes metric was calculated just over scored:<br>\n<code>y_true = tf.gather(y_true, indices=self.mask, axis=1)</code><br>\n<code>y_pred = tf.gather(y_pred, indices=self.mask, axis=1)</code></p>\n<p><strong>External data</strong><br>\n2 recent recordings from Hawaii which can be found on Xeno-Canto.<br>\nff1010bird records without birdcalls to mix with.</p>\n<p><strong>Input</strong><br>\nAs input pre-computed mel-spectrogram (mainly) or PCEN (in few cases) were used. Both calculated by Librosa with default parameters, nmels=128. So for 5s shape is 128x313.<br>\nFor mel-spectrogram primitive denoising strategy was used to reduce stationary noise over whole file. On inference same denoising approach was applied but on 5s chunk level instead whole file.</p>\n<p>In case of mel-spec, after augmentations and before mixup it was log-scaled and normed from 0 to 1.<br>\nIn case of PCEN after tiny augmentations and before mixup it was normed from 0 to 1.</p>\n<p><em>Idea to beat noisy labels a little bit (was implemented for few training attempts):</em><br>\nDrop time bins with low std (can be done for PCEN or log-melspectrogram):</p>\n<p><code>std_over_f = pcen.std(axis=0)</code><br>\n<code>pcen_f = pcen[:, std_over_f &gt; std_over_f.mean()]</code></p>\n<p>Some mask smoothing can be applied before resampling by scipy.ndimage.convolve.</p>\n<p><strong>Approaches</strong><br>\n2 different approaches were implemented:</p>\n<ol>\n<li>Multiclass classifier for all (152) species. CategoricalCrossentropy loss with label_smoothing=0.1 was used. Sigmoid is used as activation function of output layer (instead of softmax). On inference predictions were normalized from 0 to 1 (divided by max value).</li>\n<li>Multilabel classifier for chosen species. In addition to scored labels the species which have high overlap with scored (in secondary labels) or which have big amount of recordings  are chosen. Records for unchosen species are used for random mixing with initial input.<br>\ntf.nn.weighted_cross_entropy_with_logits with added label smoothing was used (weight=4.0, label_smoothing=0.1).</li>\n</ol>\n<p>ff1010bird dataset was also tried as initial input as nocall class in 1st approach and as absence of all classes in the 2nd instead of use it for mixing.</p>\n<p>Mainly the efficientnets (B0, B2, B3) were used with noisy-student weights, also seresnet101 was trained for mel-spec multiclass classifier.<br>\nInstead of simple global pooling in most attempts AutoPool over one axis followed by mean pooling was used.</p>\n<p><strong>Augmentations</strong></p>\n<ol>\n<li>Random crop of 5s window inside whole file (or restricted part in case of oversample).</li>\n<li>Time stretching and compression by resize over time axis.</li>\n<li>Random padding if file shorter than 5 s.</li>\n<li>Random low pass filtering approximately done directly on mel-spec.</li>\n<li>Mix with nobird records from ff1010bird.</li>\n<li>Mix over time: up to five consecutive 5s parts can be taken and mixed with random gains. </li>\n<li>Mixup: beta=2, alpha=3*beta. Mixing lambda was forced to be &gt;=0.5 for scored label in case of mixing scored with non-scored.</li>\n</ol>\n<p><strong>Thresholds</strong><br>\nTrained models were evaluated over whole files from validation set and best thresholds were estimated. Obtained threshold for different approaches were averaged and tuned by knowledge about amount of data per bird in dataset.<br>\nMin value of 0.25 was used for 7 rare classes and just two species (skylar and houfin) have thresholds more than 0.5.</p>\n<p><strong>Ensemble</strong><br>\nAt first average predictions are estimated per each approach (mel-spec multiclass, pcen multiclass, mel-spec mixin multilabel) and then final averaging was done with coefficients selected from LB results.</p>\n<p><strong>Post-processing</strong><br>\nFor reduce FP and find out nocall parts few checks were implemented:</p>\n<ul>\n<li>std behavior for mel-spec and pcen;</li>\n<li>mean/max relations between whole file and current chunk;</li>\n<li>check scored labels max prediction over file.</li>\n</ul>\n<h1>standwithukraine</h1>\n<h1>stoprussianaggression</h1>",
  "messages": [
    {
      "id": 1800745,
      "postDate": "2022-05-25T07:23:16.767Z",
      "content": "<p>Thanks to everybody who organized this competition. Congratulations to all who received medals.<br>\nSeems that results of this competition aren't make a lot of sense because the host solution is the best.</p>\n<p>Few words about my trying, hope you find it interesting or even useful.</p>\n<p><strong>Metrics</strong><br>\nI start with macro Fb-score which seems are suitable for case of noisy labels. I chose quite big beta - 16 to boost recall.<br>\nAfter reading some comments of hosts about the metric I switched to:<br>\n<code>M = (1/N) * (1/21) * ∑∑( (TP/(TP+FN) + TN/(TN+FP))/2 )</code></p>\n<p>For any amount of training classes metric was calculated just over scored:<br>\n<code>y_true = tf.gather(y_true, indices=self.mask, axis=1)</code><br>\n<code>y_pred = tf.gather(y_pred, indices=self.mask, axis=1)</code></p>\n<p><strong>External data</strong><br>\n2 recent recordings from Hawaii which can be found on Xeno-Canto.<br>\nff1010bird records without birdcalls to mix with.</p>\n<p><strong>Input</strong><br>\nAs input pre-computed mel-spectrogram (mainly) or PCEN (in few cases) were used. Both calculated by Librosa with default parameters, nmels=128. So for 5s shape is 128x313.<br>\nFor mel-spectrogram primitive denoising strategy was used to reduce stationary noise over whole file. On inference same denoising approach was applied but on 5s chunk level instead whole file.</p>\n<p>In case of mel-spec, after augmentations and before mixup it was log-scaled and normed from 0 to 1.<br>\nIn case of PCEN after tiny augmentations and before mixup it was normed from 0 to 1.</p>\n<p><em>Idea to beat noisy labels a little bit (was implemented for few training attempts):</em><br>\nDrop time bins with low std (can be done for PCEN or log-melspectrogram):</p>\n<p><code>std_over_f = pcen.std(axis=0)</code><br>\n<code>pcen_f = pcen[:, std_over_f &gt; std_over_f.mean()]</code></p>\n<p>Some mask smoothing can be applied before resampling by scipy.ndimage.convolve.</p>\n<p><strong>Approaches</strong><br>\n2 different approaches were implemented:</p>\n<ol>\n<li>Multiclass classifier for all (152) species. CategoricalCrossentropy loss with label_smoothing=0.1 was used. Sigmoid is used as activation function of output layer (instead of softmax). On inference predictions were normalized from 0 to 1 (divided by max value).</li>\n<li>Multilabel classifier for chosen species. In addition to scored labels the species which have high overlap with scored (in secondary labels) or which have big amount of recordings  are chosen. Records for unchosen species are used for random mixing with initial input.<br>\ntf.nn.weighted_cross_entropy_with_logits with added label smoothing was used (weight=4.0, label_smoothing=0.1).</li>\n</ol>\n<p>ff1010bird dataset was also tried as initial input as nocall class in 1st approach and as absence of all classes in the 2nd instead of use it for mixing.</p>\n<p>Mainly the efficientnets (B0, B2, B3) were used with noisy-student weights, also seresnet101 was trained for mel-spec multiclass classifier.<br>\nInstead of simple global pooling in most attempts AutoPool over one axis followed by mean pooling was used.</p>\n<p><strong>Augmentations</strong></p>\n<ol>\n<li>Random crop of 5s window inside whole file (or restricted part in case of oversample).</li>\n<li>Time stretching and compression by resize over time axis.</li>\n<li>Random padding if file shorter than 5 s.</li>\n<li>Random low pass filtering approximately done directly on mel-spec.</li>\n<li>Mix with nobird records from ff1010bird.</li>\n<li>Mix over time: up to five consecutive 5s parts can be taken and mixed with random gains. </li>\n<li>Mixup: beta=2, alpha=3*beta. Mixing lambda was forced to be &gt;=0.5 for scored label in case of mixing scored with non-scored.</li>\n</ol>\n<p><strong>Thresholds</strong><br>\nTrained models were evaluated over whole files from validation set and best thresholds were estimated. Obtained threshold for different approaches were averaged and tuned by knowledge about amount of data per bird in dataset.<br>\nMin value of 0.25 was used for 7 rare classes and just two species (skylar and houfin) have thresholds more than 0.5.</p>\n<p><strong>Ensemble</strong><br>\nAt first average predictions are estimated per each approach (mel-spec multiclass, pcen multiclass, mel-spec mixin multilabel) and then final averaging was done with coefficients selected from LB results.</p>\n<p><strong>Post-processing</strong><br>\nFor reduce FP and find out nocall parts few checks were implemented:</p>\n<ul>\n<li>std behavior for mel-spec and pcen;</li>\n<li>mean/max relations between whole file and current chunk;</li>\n<li>check scored labels max prediction over file.</li>\n</ul>\n<h1>standwithukraine</h1>\n<h1>stoprussianaggression</h1>",
      "rawMarkdown": "Thanks to everybody who organized this competition. Congratulations to all who received medals.\nSeems that results of this competition aren't make a lot of sense because the host solution is the best.\n\nFew words about my trying, hope you find it interesting or even useful.\n\n**Metrics**\nI start with macro Fb-score which seems are suitable for case of noisy labels. I chose quite big beta - 16 to boost recall.\nAfter reading some comments of hosts about the metric I switched to:\n`M = (1/N) * (1/21) * ∑∑( (TP/(TP+FN) + TN/(TN+FP))/2 )`\n\nFor any amount of training classes metric was calculated just over scored:\n`y_true = tf.gather(y_true, indices=self.mask, axis=1)`\n`y_pred = tf.gather(y_pred, indices=self.mask, axis=1)`\n\n**External data**\n2 recent recordings from Hawaii which can be found on Xeno-Canto.\nff1010bird records without birdcalls to mix with.\n\n**Input**\nAs input pre-computed mel-spectrogram (mainly) or PCEN (in few cases) were used. Both calculated by Librosa with default parameters, nmels=128. So for 5s shape is 128x313.\nFor mel-spectrogram primitive denoising strategy was used to reduce stationary noise over whole file. On inference same denoising approach was applied but on 5s chunk level instead whole file.\n\nIn case of mel-spec, after augmentations and before mixup it was log-scaled and normed from 0 to 1.\nIn case of PCEN after tiny augmentations and before mixup it was normed from 0 to 1.\n\n*Idea to beat noisy labels a little bit (was implemented for few training attempts):*\nDrop time bins with low std (can be done for PCEN or log-melspectrogram):\n\n`std_over_f = pcen.std(axis=0)`\n`pcen_f = pcen[:, std_over_f > std_over_f.mean()]`\n\nSome mask smoothing can be applied before resampling by scipy.ndimage.convolve.\n\n**Approaches**\n2 different approaches were implemented:\n1.  Multiclass classifier for all (152) species. CategoricalCrossentropy loss with label_smoothing=0.1 was used. Sigmoid is used as activation function of output layer (instead of softmax). On inference predictions were normalized from 0 to 1 (divided by max value).\n2. Multilabel classifier for chosen species. In addition to scored labels the species which have high overlap with scored (in secondary labels) or which have big amount of recordings  are chosen. Records for unchosen species are used for random mixing with initial input.\ntf.nn.weighted_cross_entropy_with_logits with added label smoothing was used (weight=4.0, label_smoothing=0.1).\n\nff1010bird dataset was also tried as initial input as nocall class in 1st approach and as absence of all classes in the 2nd instead of use it for mixing.\n\nMainly the efficientnets (B0, B2, B3) were used with noisy-student weights, also seresnet101 was trained for mel-spec multiclass classifier.\nInstead of simple global pooling in most attempts AutoPool over one axis followed by mean pooling was used.\n\n**Augmentations**\n1. Random crop of 5s window inside whole file (or restricted part in case of oversample).\n2. Time stretching and compression by resize over time axis.\n3. Random padding if file shorter than 5 s.\n4. Random low pass filtering approximately done directly on mel-spec.\n5. Mix with nobird records from ff1010bird.\n6. Mix over time: up to five consecutive 5s parts can be taken and mixed with random gains. \n7. Mixup: beta=2, alpha=3*beta. Mixing lambda was forced to be >=0.5 for scored label in case of mixing scored with non-scored.\n\n**Thresholds**\nTrained models were evaluated over whole files from validation set and best thresholds were estimated. Obtained threshold for different approaches were averaged and tuned by knowledge about amount of data per bird in dataset.\nMin value of 0.25 was used for 7 rare classes and just two species (skylar and houfin) have thresholds more than 0.5.\n\n**Ensemble**\nAt first average predictions are estimated per each approach (mel-spec multiclass, pcen multiclass, mel-spec mixin multilabel) and then final averaging was done with coefficients selected from LB results.\n\n**Post-processing**\nFor reduce FP and find out nocall parts few checks were implemented:\n- std behavior for mel-spec and pcen;\n- mean/max relations between whole file and current chunk;\n- check scored labels max prediction over file.\n\n#standwithukraine\n#stoprussianaggression",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1800745": "Thanks to everybody who organized this competition. Congratulations to all who received medals.\nSeems that results of this competition aren't make a lot of sense because the host solution is the best.\n\nFew words about my trying, hope you find it interesting or even useful.\n\n**Metrics**\nI start with macro Fb-score which seems are suitable for case of noisy labels. I chose quite big beta - 16 to boost recall.\nAfter reading some comments of hosts about the metric I switched to:\n`M = (1/N) * (1/21) * ∑∑( (TP/(TP+FN) + TN/(TN+FP))/2 )`\n\nFor any amount of training classes metric was calculated just over scored:\n`y_true = tf.gather(y_true, indices=self.mask, axis=1)`\n`y_pred = tf.gather(y_pred, indices=self.mask, axis=1)`\n\n**External data**\n2 recent recordings from Hawaii which can be found on Xeno-Canto.\nff1010bird records without birdcalls to mix with.\n\n**Input**\nAs input pre-computed mel-spectrogram (mainly) or PCEN (in few cases) were used. Both calculated by Librosa with default parameters, nmels=128. So for 5s shape is 128x313.\nFor mel-spectrogram primitive denoising strategy was used to reduce stationary noise over whole file. On inference same denoising approach was applied but on 5s chunk level instead whole file.\n\nIn case of mel-spec, after augmentations and before mixup it was log-scaled and normed from 0 to 1.\nIn case of PCEN after tiny augmentations and before mixup it was normed from 0 to 1.\n\n*Idea to beat noisy labels a little bit (was implemented for few training attempts):*\nDrop time bins with low std (can be done for PCEN or log-melspectrogram):\n\n`std_over_f = pcen.std(axis=0)`\n`pcen_f = pcen[:, std_over_f > std_over_f.mean()]`\n\nSome mask smoothing can be applied before resampling by scipy.ndimage.convolve.\n\n**Approaches**\n2 different approaches were implemented:\n1.  Multiclass classifier for all (152) species. CategoricalCrossentropy loss with label_smoothing=0.1 was used. Sigmoid is used as activation function of output layer (instead of softmax). On inference predictions were normalized from 0 to 1 (divided by max value).\n2. Multilabel classifier for chosen species. In addition to scored labels the species which have high overlap with scored (in secondary labels) or which have big amount of recordings  are chosen. Records for unchosen species are used for random mixing with initial input.\ntf.nn.weighted_cross_entropy_with_logits with added label smoothing was used (weight=4.0, label_smoothing=0.1).\n\nff1010bird dataset was also tried as initial input as nocall class in 1st approach and as absence of all classes in the 2nd instead of use it for mixing.\n\nMainly the efficientnets (B0, B2, B3) were used with noisy-student weights, also seresnet101 was trained for mel-spec multiclass classifier.\nInstead of simple global pooling in most attempts AutoPool over one axis followed by mean pooling was used.\n\n**Augmentations**\n1. Random crop of 5s window inside whole file (or restricted part in case of oversample).\n2. Time stretching and compression by resize over time axis.\n3. Random padding if file shorter than 5 s.\n4. Random low pass filtering approximately done directly on mel-spec.\n5. Mix with nobird records from ff1010bird.\n6. Mix over time: up to five consecutive 5s parts can be taken and mixed with random gains. \n7. Mixup: beta=2, alpha=3*beta. Mixing lambda was forced to be >=0.5 for scored label in case of mixing scored with non-scored.\n\n**Thresholds**\nTrained models were evaluated over whole files from validation set and best thresholds were estimated. Obtained threshold for different approaches were averaged and tuned by knowledge about amount of data per bird in dataset.\nMin value of 0.25 was used for 7 rare classes and just two species (skylar and houfin) have thresholds more than 0.5.\n\n**Ensemble**\nAt first average predictions are estimated per each approach (mel-spec multiclass, pcen multiclass, mel-spec mixin multilabel) and then final averaging was done with coefficients selected from LB results.\n\n**Post-processing**\nFor reduce FP and find out nocall parts few checks were implemented:\n- std behavior for mel-spec and pcen;\n- mean/max relations between whole file and current chunk;\n- check scored labels max prediction over file.\n\n#standwithukraine\n#stoprussianaggression"
  }
}