{
  "id": 183255,
  "title": "17th Place Solution : file-level post-process",
  "url": "/competitions/birdsong-recognition/writeups/17th-place-solution-file-level-post-process",
  "author_name": "",
  "post_date": "2020-09-17T16:08:20.357Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p><a href=\"https://github.com/xins-yao/Kaggle_Birdcall_17th_solution\" target=\"_blank\">https://github.com/xins-yao/Kaggle_Birdcall_17th_solution</a></p>\n<h2>Feature Engineering</h2>\n<ul>\n<li>Log MEL Spectrogram</li>\n</ul>\n<pre><code>sr = 32000\nfmin = 20\nfmax = sr // 2\n\nn_channel = 128\nn_fft = 2048\nhop_length = 512\nwin_length = n_fft\n</code></pre>\n<h2>Augment</h2>\n<ul>\n<li>Random Clip<br>\nrandomly cut 5s clip from audio and only keep clips with SNR higher than 1e-3</li>\n</ul>\n<pre><code>  def signal_noise_ratio(spec):\n      spec = spec.copy()\n\n      col_median = np.median(spec, axis=0, keepdims=True)\n      row_median = np.median(spec, axis=1, keepdims=True)\n\n      spec[spec &lt; row_median * 1.25] = 0.0\n      spec[spec &lt; col_median * 1.15] = 0.0\n      spec[spec &gt; 0] = 1.0\n\n      spec = cv2.medianBlur(spec, 3)\n      spec = cv2.morphologyEx(spec, cv2.MORPH_CLOSE, np.ones((3, 3), np.float32))\n\n      spec_sum = spec.sum()\n      try:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1] * spec.shape[2])\n      except:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1])\n\n      return snr\n</code></pre>\n<ul>\n<li>MixUp<br>\nmixup over LogMelSpec with beta(8.0, 8.0) distribution<br>\nbeta(0.4, 0.4) and beta(1.0, 1.0) will raise more TruePositive but lead to much more FalsePositive</li>\n<li>Noise<br>\nadd up to 4 noises with independent probabilities and scales in waveform<br>\nnoises were extracted from training samples</li>\n</ul>\n<pre><code>  def signal_noise_split(audio):\n    S, _ = spectrum._spectrogram(y=audio, power=1.0, n_fft=2048, hop_length=512, win_length=2048)\n\n    col_median = np.median(S, axis=0, keepdims=True)\n    row_median = np.median(S, axis=1, keepdims=True)\n    S[S &lt; row_median * 3] = 0.0\n    S[S &lt; col_median * 3] = 0.0\n    S[S &gt; 0] = 1\n\n    S = binary_erosion(S, structure=np.ones((4, 4)))\n    S = binary_dilation(S, structure=np.ones((4, 4)))\n\n    indicator = S.any(axis=0)\n    indicator = binary_dilation(indicator, structure=np.ones(4), iterations=2)\n\n    mask = np.repeat(indicator, hop_length)\n    mask = binary_dilation(mask, structure=np.ones(win_length - hop_length), origin=-(win_length - hop_length)//2)\n    mask = mask[:len(audio)]\n    signal = audio[mask]\n    noise = audio[~mask]\n    return signal, noise\n</code></pre>\n<h2>Model</h2>\n<ul>\n<li>CNN<br>\n9-layer CNN<br>\naverage pooling over frequency before max pooling over time within each ConvBlock2D<br>\nSqueezeExcitationBlock within each ConvBlock2D<br>\npixel shuffle: (n_channel, n_freq, n_time) -&gt; (n_channel * 2, n_freq / 2, n_time)</li>\n<li>CRNN<br>\n2-layer bidirectional GRU after 9-layer CNN</li>\n<li>CNN + Transformer Encoder<br>\nEncoder with 8-AttentionHead after 9-layer CNN</li>\n</ul>\n<h2>Trainng</h2>\n<ul>\n<li>Label Smooth: 0.05 alpha</li>\n<li>Balance Sampler: randomly select up to 150 samples of each bird</li>\n<li>Stratified 5Fold based on ebird_code</li>\n<li>Loss Function: BCEWithLogitsLoss</li>\n<li>Optimizer: Adam(lr=1e-3)</li>\n<li>Scheduler: CosineAnnealingLR(Tmax=10)</li>\n</ul>\n<h2>Post-Process</h2>\n<p>if model gets a confident prediction of any bird, then lower threshold for this bird in the same audio file</p>\n<ul>\n<li>use thr_median as initial threshold</li>\n<li>use thr_high for confident prediction</li>\n<li>if any bird with probability higher than thr_high in any clip, lower threshold to thr_low for this specific bird in the same audio file</li>\n</ul>",
  "messages": [
    {
      "id": "1012369",
      "postDate": "09/16/2020 03:47:57",
      "content": "<p><a href=\"https://github.com/xins-yao/Kaggle_Birdcall_17th_solution\" target=\"_blank\">https://github.com/xins-yao/Kaggle_Birdcall_17th_solution</a></p>\n<h2>Feature Engineering</h2>\n<ul>\n<li>Log MEL Spectrogram</li>\n</ul>\n<pre><code>sr = 32000\nfmin = 20\nfmax = sr // 2\n\nn_channel = 128\nn_fft = 2048\nhop_length = 512\nwin_length = n_fft\n</code></pre>\n<h2>Augment</h2>\n<ul>\n<li>Random Clip<br>\nrandomly cut 5s clip from audio and only keep clips with SNR higher than 1e-3</li>\n</ul>\n<pre><code>  def signal_noise_ratio(spec):\n      spec = spec.copy()\n\n      col_median = np.median(spec, axis=0, keepdims=True)\n      row_median = np.median(spec, axis=1, keepdims=True)\n\n      spec[spec &lt; row_median * 1.25] = 0.0\n      spec[spec &lt; col_median * 1.15] = 0.0\n      spec[spec &gt; 0] = 1.0\n\n      spec = cv2.medianBlur(spec, 3)\n      spec = cv2.morphologyEx(spec, cv2.MORPH_CLOSE, np.ones((3, 3), np.float32))\n\n      spec_sum = spec.sum()\n      try:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1] * spec.shape[2])\n      except:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1])\n\n      return snr\n</code></pre>\n<ul>\n<li>MixUp<br>\nmixup over LogMelSpec with beta(8.0, 8.0) distribution<br>\nbeta(0.4, 0.4) and beta(1.0, 1.0) will raise more TruePositive but lead to much more FalsePositive</li>\n<li>Noise<br>\nadd up to 4 noises with independent probabilities and scales in waveform<br>\nnoises were extracted from training samples</li>\n</ul>\n<pre><code>  def signal_noise_split(audio):\n    S, _ = spectrum._spectrogram(y=audio, power=1.0, n_fft=2048, hop_length=512, win_length=2048)\n\n    col_median = np.median(S, axis=0, keepdims=True)\n    row_median = np.median(S, axis=1, keepdims=True)\n    S[S &lt; row_median * 3] = 0.0\n    S[S &lt; col_median * 3] = 0.0\n    S[S &gt; 0] = 1\n\n    S = binary_erosion(S, structure=np.ones((4, 4)))\n    S = binary_dilation(S, structure=np.ones((4, 4)))\n\n    indicator = S.any(axis=0)\n    indicator = binary_dilation(indicator, structure=np.ones(4), iterations=2)\n\n    mask = np.repeat(indicator, hop_length)\n    mask = binary_dilation(mask, structure=np.ones(win_length - hop_length), origin=-(win_length - hop_length)//2)\n    mask = mask[:len(audio)]\n    signal = audio[mask]\n    noise = audio[~mask]\n    return signal, noise\n</code></pre>\n<h2>Model</h2>\n<ul>\n<li>CNN<br>\n9-layer CNN<br>\naverage pooling over frequency before max pooling over time within each ConvBlock2D<br>\nSqueezeExcitationBlock within each ConvBlock2D<br>\npixel shuffle: (n_channel, n_freq, n_time) -&gt; (n_channel * 2, n_freq / 2, n_time)</li>\n<li>CRNN<br>\n2-layer bidirectional GRU after 9-layer CNN</li>\n<li>CNN + Transformer Encoder<br>\nEncoder with 8-AttentionHead after 9-layer CNN</li>\n</ul>\n<h2>Trainng</h2>\n<ul>\n<li>Label Smooth: 0.05 alpha</li>\n<li>Balance Sampler: randomly select up to 150 samples of each bird</li>\n<li>Stratified 5Fold based on ebird_code</li>\n<li>Loss Function: BCEWithLogitsLoss</li>\n<li>Optimizer: Adam(lr=1e-3)</li>\n<li>Scheduler: CosineAnnealingLR(Tmax=10)</li>\n</ul>\n<h2>Post-Process</h2>\n<p>if model gets a confident prediction of any bird, then lower threshold for this bird in the same audio file</p>\n<ul>\n<li>use thr_median as initial threshold</li>\n<li>use thr_high for confident prediction</li>\n<li>if any bird with probability higher than thr_high in any clip, lower threshold to thr_low for this specific bird in the same audio file</li>\n</ul>",
      "rawMarkdown": "https://github.com/xins-yao/Kaggle_Birdcall_17th_solution\n## Feature Engineering\n- Log MEL Spectrogram\n```\nsr = 32000\nfmin = 20\nfmax = sr // 2\n\nn_channel = 128\nn_fft = 2048\nhop_length = 512\nwin_length = n_fft\n```\n\n## Augment\n- Random Clip\n randomly cut 5s clip from audio and only keep clips with SNR higher than 1e-3\n  ```\n  def signal_noise_ratio(spec):\n      spec = spec.copy()\n\n      col_median = np.median(spec, axis=0, keepdims=True)\n      row_median = np.median(spec, axis=1, keepdims=True)\n\n      spec[spec < row_median * 1.25] = 0.0\n      spec[spec < col_median * 1.15] = 0.0\n      spec[spec > 0] = 1.0\n\n      spec = cv2.medianBlur(spec, 3)\n      spec = cv2.morphologyEx(spec, cv2.MORPH_CLOSE, np.ones((3, 3), np.float32))\n\n      spec_sum = spec.sum()\n      try:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1] * spec.shape[2])\n      except:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1])\n\n      return snr\n  ```\n- MixUp\n mixup over LogMelSpec with beta(8.0, 8.0) distribution\n beta(0.4, 0.4) and beta(1.0, 1.0) will raise more TruePositive but lead to much more FalsePositive\n- Noise\n add up to 4 noises with independent probabilities and scales in waveform\n noises were extracted from training samples\n  ```\n  def signal_noise_split(audio):\n    S, _ = spectrum._spectrogram(y=audio, power=1.0, n_fft=2048, hop_length=512, win_length=2048)\n    \n    col_median = np.median(S, axis=0, keepdims=True)\n    row_median = np.median(S, axis=1, keepdims=True)\n    S[S < row_median * 3] = 0.0\n    S[S < col_median * 3] = 0.0\n    S[S > 0] = 1\n    \n    S = binary_erosion(S, structure=np.ones((4, 4)))\n    S = binary_dilation(S, structure=np.ones((4, 4)))\n    \n    indicator = S.any(axis=0)\n    indicator = binary_dilation(indicator, structure=np.ones(4), iterations=2)\n    \n    mask = np.repeat(indicator, hop_length)\n    mask = binary_dilation(mask, structure=np.ones(win_length - hop_length), origin=-(win_length - hop_length)//2)\n    mask = mask[:len(audio)]\n    signal = audio[mask]\n    noise = audio[~mask]\n    return signal, noise\n  ```\n## Model\n- CNN\n 9-layer CNN\n average pooling over frequency before max pooling over time within each ConvBlock2D\n SqueezeExcitationBlock within each ConvBlock2D\n pixel shuffle: (n_channel, n_freq, n_time) -> (n_channel * 2, n_freq / 2, n_time)\n- CRNN\n 2-layer bidirectional GRU after 9-layer CNN\n- CNN + Transformer Encoder\n Encoder with 8-AttentionHead after 9-layer CNN\n\n## Trainng\n- Label Smooth: 0.05 alpha\n- Balance Sampler: randomly select up to 150 samples of each bird\n- Stratified 5Fold based on ebird_code\n- Loss Function: BCEWithLogitsLoss\n- Optimizer: Adam(lr=1e-3)\n- Scheduler: CosineAnnealingLR(Tmax=10)\n\n## Post-Process\n\nif model gets a confident prediction of any bird, then lower threshold for this bird in the same audio file\n- use thr_median as initial threshold\n- use thr_high for confident prediction\n- if any bird with probability higher than thr_high in any clip, lower threshold to thr_low for this specific bird in the same audio file",
      "votes": null
    },
    {
      "id": "1012739",
      "postDate": "09/16/2020 09:01:36",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "rawMarkdown": "Thank you for sharing your solution, congratulation👍",
      "votes": null
    },
    {
      "id": "1012758",
      "postDate": "09/16/2020 09:17:58",
      "content": "<p>Congrats on the result. Your file level post processing is a great idea, thanks for sharing!</p>",
      "rawMarkdown": "Congrats on the result. Your file level post processing is a great idea, thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012739,
      "author_name": "",
      "author_url": "",
      "post_date": "09/16/2020 09:01:36",
      "content": "<p>Thank you for sharing your solution, congratulation👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1012758,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/16/2020 09:17:58",
      "content": "<p>Congrats on the result. Your file level post processing is a great idea, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012369": "https://github.com/xins-yao/Kaggle_Birdcall_17th_solution\n## Feature Engineering\n- Log MEL Spectrogram\n```\nsr = 32000\nfmin = 20\nfmax = sr // 2\n\nn_channel = 128\nn_fft = 2048\nhop_length = 512\nwin_length = n_fft\n```\n\n## Augment\n- Random Clip\n randomly cut 5s clip from audio and only keep clips with SNR higher than 1e-3\n  ```\n  def signal_noise_ratio(spec):\n      spec = spec.copy()\n\n      col_median = np.median(spec, axis=0, keepdims=True)\n      row_median = np.median(spec, axis=1, keepdims=True)\n\n      spec[spec < row_median * 1.25] = 0.0\n      spec[spec < col_median * 1.15] = 0.0\n      spec[spec > 0] = 1.0\n\n      spec = cv2.medianBlur(spec, 3)\n      spec = cv2.morphologyEx(spec, cv2.MORPH_CLOSE, np.ones((3, 3), np.float32))\n\n      spec_sum = spec.sum()\n      try:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1] * spec.shape[2])\n      except:\n          snr = spec_sum / (spec.shape[0] * spec.shape[1])\n\n      return snr\n  ```\n- MixUp\n mixup over LogMelSpec with beta(8.0, 8.0) distribution\n beta(0.4, 0.4) and beta(1.0, 1.0) will raise more TruePositive but lead to much more FalsePositive\n- Noise\n add up to 4 noises with independent probabilities and scales in waveform\n noises were extracted from training samples\n  ```\n  def signal_noise_split(audio):\n    S, _ = spectrum._spectrogram(y=audio, power=1.0, n_fft=2048, hop_length=512, win_length=2048)\n    \n    col_median = np.median(S, axis=0, keepdims=True)\n    row_median = np.median(S, axis=1, keepdims=True)\n    S[S < row_median * 3] = 0.0\n    S[S < col_median * 3] = 0.0\n    S[S > 0] = 1\n    \n    S = binary_erosion(S, structure=np.ones((4, 4)))\n    S = binary_dilation(S, structure=np.ones((4, 4)))\n    \n    indicator = S.any(axis=0)\n    indicator = binary_dilation(indicator, structure=np.ones(4), iterations=2)\n    \n    mask = np.repeat(indicator, hop_length)\n    mask = binary_dilation(mask, structure=np.ones(win_length - hop_length), origin=-(win_length - hop_length)//2)\n    mask = mask[:len(audio)]\n    signal = audio[mask]\n    noise = audio[~mask]\n    return signal, noise\n  ```\n## Model\n- CNN\n 9-layer CNN\n average pooling over frequency before max pooling over time within each ConvBlock2D\n SqueezeExcitationBlock within each ConvBlock2D\n pixel shuffle: (n_channel, n_freq, n_time) -> (n_channel * 2, n_freq / 2, n_time)\n- CRNN\n 2-layer bidirectional GRU after 9-layer CNN\n- CNN + Transformer Encoder\n Encoder with 8-AttentionHead after 9-layer CNN\n\n## Trainng\n- Label Smooth: 0.05 alpha\n- Balance Sampler: randomly select up to 150 samples of each bird\n- Stratified 5Fold based on ebird_code\n- Loss Function: BCEWithLogitsLoss\n- Optimizer: Adam(lr=1e-3)\n- Scheduler: CosineAnnealingLR(Tmax=10)\n\n## Post-Process\n\nif model gets a confident prediction of any bird, then lower threshold for this bird in the same audio file\n- use thr_median as initial threshold\n- use thr_high for confident prediction\n- if any bird with probability higher than thr_high in any clip, lower threshold to thr_low for this specific bird in the same audio file",
    "1012739": "Thank you for sharing your solution, congratulation👍",
    "1012758": "Congrats on the result. Your file level post processing is a great idea, thanks for sharing!"
  },
  "source": "meta"
}