{
  "id": 243332,
  "title": "49 place solution - meh models, +0.04-0.05 private LB postprocessing",
  "url": "/competitions/birdclef-2021/discussion/243332",
  "author_name": "fffrrt",
  "post_date": "2021-06-02T05:29:56.832000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>TL;DR</p>\n<ul>\n<li>6 models ensembling (5-fold each):<ul>\n<li>EfficientNet B2 pcen 128x500, same with mel, trained on random 5 second slices</li>\n<li>Same two models, trained on 5 second slices from first 20 seconds of the short audio files</li>\n<li>Rexnet200 pcen 128x400, same with mel, trained on random 5 second slices</li></ul></li>\n<li>Normalizing model predictions so any single model would not drown other models when averaging</li>\n<li>Post-processing based on false positive/true positive ratio of the birds</li>\n<li>Post-processing based on time and distance</li>\n<li>Safe threshold choosing</li>\n</ul>\n<p><strong>Data preparation</strong>:</p>\n<p>Spectrograms from all \"short\" train audio files were computed and saved on disk. (~150 GB each).<br>\nPcen spectrograms were used with fmin=100, fmax=12000, fft = 1024, hop = 320.<br>\nMel spectrograms were used with fmin=100, fmax=16000, fft = 1024, hop = 320.<br>\nAdditionally, mel spectrograms were computed several times with different power setting (in cycle from 1.0, 1.2, … 3.0), until average brightness of resulting spec was &gt;20, and that result was chosen. That was done to avoid very dark and very bright results.</p>\n<p><strong>Training</strong>:</p>\n<p>AdamW on default mostly, for ~15 epoch with LR 0.001 for first ~10-12, 0.00015 for next ~3 and 0.00003 for the last 1-2. SGD did not work at all.</p>\n<p>EfficientNet B2 was the best for me, Rexnet200 trailed ~0.015 CV behind and pcen specs were superior to mels by ~0.001. ResNeSt was significantly behind for some reason. Did not try other heavy weight nets because of computing power limitations.</p>\n<p>Due to weak GPU (RTX 2060 Super) and large amount of data - even with mixed precision training i could realistically train at most one 5-fold a day, and so experimenting was limited. At the end i do not think that i managed to beat kkiller's public notebook with any of my models, but some were in 0.01 range from it.</p>\n<p><strong>Development</strong>:</p>\n<p>First of all i was confused when i read that bagging improved score from 0.63 to 0.73. In my submissions simple averaging sometimes even reduced the score compared to individual models!</p>\n<p>After some investigation i found that logits on which i averaged were extremely different between models (probably due to different training schedule), and, probably, any subtle difference between classifications were overshadowed by that scale difference.</p>\n<p>Normalizing models predictions improved scores by ~0.02, but still nowhere near such giant jumps, just more in line of usual ensembling.</p>\n<p>Then i started predicting with 1 second steps and average all predictions that were done for some second, and taking max of five 1-second results in the soundscape. This way bird/nobird transitions were more smooth, and any possible \"bird call got sliced in two and failed to be predicted\" situations disappeared. That improved score by another ~0.01.</p>\n<p>Then i spent ~10 days trying to improve my models with pseudolabeling, hand labeling, location/time concats to features, different losses, different augments, different training schedules, different type of models, etc etc. Nothing worked good enough. Some models went into final ensemble, but at the end my strongest single model was my first one, EfficientNet B2 128x500 on pcen specs with barely any augments. This extra models for ensemble added ~0.01 to the final score.</p>\n<p>At this time there was a large 0.74+ group and a large 0.72- group on the leaderboard. I decided to find the difference between them and started to try post-processing.</p>\n<p><strong>Post-processing</strong>:</p>\n<p>Then, when looking at submission (on train soundscapes), i noticed large flocks of grhowl. From there post-processing based on false_positive / true_positive ratio was born. grhowl was the sole, giant, woo-hoo-ing offender. Any sufficiently noisy record was grhowl. Rest of the birds were significantly more behaved, and even if i slightly reduced some birds in the end, it made almost no dfiference to just removing grhowl. That improved score by ~0.015.</p>\n<p>Next were experiments with distance and time. First, most simple attempt was finding minimum distance from short train audios to recording locations and predicting zero for any bird above some threshold. That kinda worked, but improvement was minimal (~0.005)</p>\n<p>Next i decided to take distance of n% closest samples, instead of the closest one - for example, for 500 records birds i took low ~2%, and kept the distance of 10th closest sample - to avoid outliers. This worked significantly better, and improved score by another ~0.015.</p>\n<p>Linear threshold is not based on something real, though. In real life birds do not hit the edge of the map and just stop there. The next iteration of post-processing used gradually increasing probability from ~7 degrees distance to ~2 degrees distance, and it worked better - another ~0.01.</p>\n<p>At the end, time was added to the same code as third dimension (with appropriate scale factor), and all settings were tuned on train soundscapes. Because amount of data points was not very big, i was forced to look at low ~6% of the samples when determining \"distance\" and increase said \"distance\" by quite a bit. It still improved score by another ~0.005.</p>\n<p><strong>Threshold choosing</strong>:</p>\n<p>Train soundscapes and public test set had significantly different optimal thresholds. That alerted me to the possibility of large threshold swings on the full private set. I estimated the final optimal threshold would probably be between them, but slightly to the train soundscapes side, and decided to play it safe. Two different thresholds were chosen - one to intentionally undershoot by a bit, one to intentionally overshoot by a bit - but still at the points where validation scores were relatively flat and only slightly lower than optimal one.</p>\n<p>I expected to lose ~0.005-0.01 from that decision in exchange for relative safety from shakeups, but conservative high threshold actually landed on the exact optimal spot of the private LB data.</p>\n<p><strong>Conclusion</strong>:</p>\n<p>That was a fun month and a fun journey from 0.54 private LB to 0.63 private LB. I did not spend a lot of time on this competiton compared to Rainforest ones, but the result was as good.</p>\n<p>Final version of postprocessing (grhowl is on by default, it might be just my models quirk) - <a href=\"https://www.kaggle.com/fffrrt/location-and-time-postprocessing-birdclef-2021\" target=\"_blank\">https://www.kaggle.com/fffrrt/location-and-time-postprocessing-birdclef-2021</a></p>\n<p>My best submission scored 0.63/0.68 with it, and 0.59/0.62 without it (with the same threshold selection process).</p>",
  "messages": [
    {
      "id": 1332388,
      "postDate": "2021-06-02T05:29:56.833Z",
      "content": "<p>TL;DR</p>\n<ul>\n<li>6 models ensembling (5-fold each):<ul>\n<li>EfficientNet B2 pcen 128x500, same with mel, trained on random 5 second slices</li>\n<li>Same two models, trained on 5 second slices from first 20 seconds of the short audio files</li>\n<li>Rexnet200 pcen 128x400, same with mel, trained on random 5 second slices</li></ul></li>\n<li>Normalizing model predictions so any single model would not drown other models when averaging</li>\n<li>Post-processing based on false positive/true positive ratio of the birds</li>\n<li>Post-processing based on time and distance</li>\n<li>Safe threshold choosing</li>\n</ul>\n<p><strong>Data preparation</strong>:</p>\n<p>Spectrograms from all \"short\" train audio files were computed and saved on disk. (~150 GB each).<br>\nPcen spectrograms were used with fmin=100, fmax=12000, fft = 1024, hop = 320.<br>\nMel spectrograms were used with fmin=100, fmax=16000, fft = 1024, hop = 320.<br>\nAdditionally, mel spectrograms were computed several times with different power setting (in cycle from 1.0, 1.2, … 3.0), until average brightness of resulting spec was &gt;20, and that result was chosen. That was done to avoid very dark and very bright results.</p>\n<p><strong>Training</strong>:</p>\n<p>AdamW on default mostly, for ~15 epoch with LR 0.001 for first ~10-12, 0.00015 for next ~3 and 0.00003 for the last 1-2. SGD did not work at all.</p>\n<p>EfficientNet B2 was the best for me, Rexnet200 trailed ~0.015 CV behind and pcen specs were superior to mels by ~0.001. ResNeSt was significantly behind for some reason. Did not try other heavy weight nets because of computing power limitations.</p>\n<p>Due to weak GPU (RTX 2060 Super) and large amount of data - even with mixed precision training i could realistically train at most one 5-fold a day, and so experimenting was limited. At the end i do not think that i managed to beat kkiller's public notebook with any of my models, but some were in 0.01 range from it.</p>\n<p><strong>Development</strong>:</p>\n<p>First of all i was confused when i read that bagging improved score from 0.63 to 0.73. In my submissions simple averaging sometimes even reduced the score compared to individual models!</p>\n<p>After some investigation i found that logits on which i averaged were extremely different between models (probably due to different training schedule), and, probably, any subtle difference between classifications were overshadowed by that scale difference.</p>\n<p>Normalizing models predictions improved scores by ~0.02, but still nowhere near such giant jumps, just more in line of usual ensembling.</p>\n<p>Then i started predicting with 1 second steps and average all predictions that were done for some second, and taking max of five 1-second results in the soundscape. This way bird/nobird transitions were more smooth, and any possible \"bird call got sliced in two and failed to be predicted\" situations disappeared. That improved score by another ~0.01.</p>\n<p>Then i spent ~10 days trying to improve my models with pseudolabeling, hand labeling, location/time concats to features, different losses, different augments, different training schedules, different type of models, etc etc. Nothing worked good enough. Some models went into final ensemble, but at the end my strongest single model was my first one, EfficientNet B2 128x500 on pcen specs with barely any augments. This extra models for ensemble added ~0.01 to the final score.</p>\n<p>At this time there was a large 0.74+ group and a large 0.72- group on the leaderboard. I decided to find the difference between them and started to try post-processing.</p>\n<p><strong>Post-processing</strong>:</p>\n<p>Then, when looking at submission (on train soundscapes), i noticed large flocks of grhowl. From there post-processing based on false_positive / true_positive ratio was born. grhowl was the sole, giant, woo-hoo-ing offender. Any sufficiently noisy record was grhowl. Rest of the birds were significantly more behaved, and even if i slightly reduced some birds in the end, it made almost no dfiference to just removing grhowl. That improved score by ~0.015.</p>\n<p>Next were experiments with distance and time. First, most simple attempt was finding minimum distance from short train audios to recording locations and predicting zero for any bird above some threshold. That kinda worked, but improvement was minimal (~0.005)</p>\n<p>Next i decided to take distance of n% closest samples, instead of the closest one - for example, for 500 records birds i took low ~2%, and kept the distance of 10th closest sample - to avoid outliers. This worked significantly better, and improved score by another ~0.015.</p>\n<p>Linear threshold is not based on something real, though. In real life birds do not hit the edge of the map and just stop there. The next iteration of post-processing used gradually increasing probability from ~7 degrees distance to ~2 degrees distance, and it worked better - another ~0.01.</p>\n<p>At the end, time was added to the same code as third dimension (with appropriate scale factor), and all settings were tuned on train soundscapes. Because amount of data points was not very big, i was forced to look at low ~6% of the samples when determining \"distance\" and increase said \"distance\" by quite a bit. It still improved score by another ~0.005.</p>\n<p><strong>Threshold choosing</strong>:</p>\n<p>Train soundscapes and public test set had significantly different optimal thresholds. That alerted me to the possibility of large threshold swings on the full private set. I estimated the final optimal threshold would probably be between them, but slightly to the train soundscapes side, and decided to play it safe. Two different thresholds were chosen - one to intentionally undershoot by a bit, one to intentionally overshoot by a bit - but still at the points where validation scores were relatively flat and only slightly lower than optimal one.</p>\n<p>I expected to lose ~0.005-0.01 from that decision in exchange for relative safety from shakeups, but conservative high threshold actually landed on the exact optimal spot of the private LB data.</p>\n<p><strong>Conclusion</strong>:</p>\n<p>That was a fun month and a fun journey from 0.54 private LB to 0.63 private LB. I did not spend a lot of time on this competiton compared to Rainforest ones, but the result was as good.</p>\n<p>Final version of postprocessing (grhowl is on by default, it might be just my models quirk) - <a href=\"https://www.kaggle.com/fffrrt/location-and-time-postprocessing-birdclef-2021\" target=\"_blank\">https://www.kaggle.com/fffrrt/location-and-time-postprocessing-birdclef-2021</a></p>\n<p>My best submission scored 0.63/0.68 with it, and 0.59/0.62 without it (with the same threshold selection process).</p>",
      "rawMarkdown": "TL;DR\n\n* 6 models ensembling (5-fold each):\n   - EfficientNet B2 pcen 128x500, same with mel, trained on random 5 second slices\n   - Same two models, trained on 5 second slices from first 20 seconds of the short audio files\n   - Rexnet200 pcen 128x400, same with mel, trained on random 5 second slices\n* Normalizing model predictions so any single model would not drown other models when averaging\n* Post-processing based on false positive/true positive ratio of the birds\n* Post-processing based on time and distance\n* Safe threshold choosing\n\n**Data preparation**:\n\nSpectrograms from all \"short\" train audio files were computed and saved on disk. (~150 GB each).\nPcen spectrograms were used with fmin=100, fmax=12000, fft = 1024, hop = 320.\nMel spectrograms were used with fmin=100, fmax=16000, fft = 1024, hop = 320.\nAdditionally, mel spectrograms were computed several times with different power setting (in cycle from 1.0, 1.2, ... 3.0), until average brightness of resulting spec was >20, and that result was chosen. That was done to avoid very dark and very bright results.\n\n**Training**:\n\nAdamW on default mostly, for ~15 epoch with LR 0.001 for first ~10-12, 0.00015 for next ~3 and 0.00003 for the last 1-2. SGD did not work at all.\n\nEfficientNet B2 was the best for me, Rexnet200 trailed ~0.015 CV behind and pcen specs were superior to mels by ~0.001. ResNeSt was significantly behind for some reason. Did not try other heavy weight nets because of computing power limitations.\n\nDue to weak GPU (RTX 2060 Super) and large amount of data - even with mixed precision training i could realistically train at most one 5-fold a day, and so experimenting was limited. At the end i do not think that i managed to beat kkiller's public notebook with any of my models, but some were in 0.01 range from it.\n\n**Development**:\n\nFirst of all i was confused when i read that bagging improved score from 0.63 to 0.73. In my submissions simple averaging sometimes even reduced the score compared to individual models!\n\nAfter some investigation i found that logits on which i averaged were extremely different between models (probably due to different training schedule), and, probably, any subtle difference between classifications were overshadowed by that scale difference.\n\nNormalizing models predictions improved scores by ~0.02, but still nowhere near such giant jumps, just more in line of usual ensembling.\n\nThen i started predicting with 1 second steps and average all predictions that were done for some second, and taking max of five 1-second results in the soundscape. This way bird/nobird transitions were more smooth, and any possible \"bird call got sliced in two and failed to be predicted\" situations disappeared. That improved score by another ~0.01.\n\nThen i spent ~10 days trying to improve my models with pseudolabeling, hand labeling, location/time concats to features, different losses, different augments, different training schedules, different type of models, etc etc. Nothing worked good enough. Some models went into final ensemble, but at the end my strongest single model was my first one, EfficientNet B2 128x500 on pcen specs with barely any augments. This extra models for ensemble added ~0.01 to the final score.\n\nAt this time there was a large 0.74+ group and a large 0.72- group on the leaderboard. I decided to find the difference between them and started to try post-processing.\n\n**Post-processing**:\n\nThen, when looking at submission (on train soundscapes), i noticed large flocks of grhowl. From there post-processing based on false_positive / true_positive ratio was born. grhowl was the sole, giant, woo-hoo-ing offender. Any sufficiently noisy record was grhowl. Rest of the birds were significantly more behaved, and even if i slightly reduced some birds in the end, it made almost no dfiference to just removing grhowl. That improved score by ~0.015.\n\nNext were experiments with distance and time. First, most simple attempt was finding minimum distance from short train audios to recording locations and predicting zero for any bird above some threshold. That kinda worked, but improvement was minimal (~0.005)\n\nNext i decided to take distance of n% closest samples, instead of the closest one - for example, for 500 records birds i took low ~2%, and kept the distance of 10th closest sample - to avoid outliers. This worked significantly better, and improved score by another ~0.015.\n\nLinear threshold is not based on something real, though. In real life birds do not hit the edge of the map and just stop there. The next iteration of post-processing used gradually increasing probability from ~7 degrees distance to ~2 degrees distance, and it worked better - another ~0.01.\n\nAt the end, time was added to the same code as third dimension (with appropriate scale factor), and all settings were tuned on train soundscapes. Because amount of data points was not very big, i was forced to look at low ~6% of the samples when determining \"distance\" and increase said \"distance\" by quite a bit. It still improved score by another ~0.005.\n\n**Threshold choosing**:\n\nTrain soundscapes and public test set had significantly different optimal thresholds. That alerted me to the possibility of large threshold swings on the full private set. I estimated the final optimal threshold would probably be between them, but slightly to the train soundscapes side, and decided to play it safe. Two different thresholds were chosen - one to intentionally undershoot by a bit, one to intentionally overshoot by a bit - but still at the points where validation scores were relatively flat and only slightly lower than optimal one.\n\nI expected to lose ~0.005-0.01 from that decision in exchange for relative safety from shakeups, but conservative high threshold actually landed on the exact optimal spot of the private LB data.\n\n**Conclusion**:\n\nThat was a fun month and a fun journey from 0.54 private LB to 0.63 private LB. I did not spend a lot of time on this competiton compared to Rainforest ones, but the result was as good.\n\nFinal version of postprocessing (grhowl is on by default, it might be just my models quirk) - https://www.kaggle.com/fffrrt/location-and-time-postprocessing-birdclef-2021\n\nMy best submission scored 0.63/0.68 with it, and 0.59/0.62 without it (with the same threshold selection process).",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1332388": "TL;DR\n\n* 6 models ensembling (5-fold each):\n   - EfficientNet B2 pcen 128x500, same with mel, trained on random 5 second slices\n   - Same two models, trained on 5 second slices from first 20 seconds of the short audio files\n   - Rexnet200 pcen 128x400, same with mel, trained on random 5 second slices\n* Normalizing model predictions so any single model would not drown other models when averaging\n* Post-processing based on false positive/true positive ratio of the birds\n* Post-processing based on time and distance\n* Safe threshold choosing\n\n**Data preparation**:\n\nSpectrograms from all \"short\" train audio files were computed and saved on disk. (~150 GB each).\nPcen spectrograms were used with fmin=100, fmax=12000, fft = 1024, hop = 320.\nMel spectrograms were used with fmin=100, fmax=16000, fft = 1024, hop = 320.\nAdditionally, mel spectrograms were computed several times with different power setting (in cycle from 1.0, 1.2, ... 3.0), until average brightness of resulting spec was >20, and that result was chosen. That was done to avoid very dark and very bright results.\n\n**Training**:\n\nAdamW on default mostly, for ~15 epoch with LR 0.001 for first ~10-12, 0.00015 for next ~3 and 0.00003 for the last 1-2. SGD did not work at all.\n\nEfficientNet B2 was the best for me, Rexnet200 trailed ~0.015 CV behind and pcen specs were superior to mels by ~0.001. ResNeSt was significantly behind for some reason. Did not try other heavy weight nets because of computing power limitations.\n\nDue to weak GPU (RTX 2060 Super) and large amount of data - even with mixed precision training i could realistically train at most one 5-fold a day, and so experimenting was limited. At the end i do not think that i managed to beat kkiller's public notebook with any of my models, but some were in 0.01 range from it.\n\n**Development**:\n\nFirst of all i was confused when i read that bagging improved score from 0.63 to 0.73. In my submissions simple averaging sometimes even reduced the score compared to individual models!\n\nAfter some investigation i found that logits on which i averaged were extremely different between models (probably due to different training schedule), and, probably, any subtle difference between classifications were overshadowed by that scale difference.\n\nNormalizing models predictions improved scores by ~0.02, but still nowhere near such giant jumps, just more in line of usual ensembling.\n\nThen i started predicting with 1 second steps and average all predictions that were done for some second, and taking max of five 1-second results in the soundscape. This way bird/nobird transitions were more smooth, and any possible \"bird call got sliced in two and failed to be predicted\" situations disappeared. That improved score by another ~0.01.\n\nThen i spent ~10 days trying to improve my models with pseudolabeling, hand labeling, location/time concats to features, different losses, different augments, different training schedules, different type of models, etc etc. Nothing worked good enough. Some models went into final ensemble, but at the end my strongest single model was my first one, EfficientNet B2 128x500 on pcen specs with barely any augments. This extra models for ensemble added ~0.01 to the final score.\n\nAt this time there was a large 0.74+ group and a large 0.72- group on the leaderboard. I decided to find the difference between them and started to try post-processing.\n\n**Post-processing**:\n\nThen, when looking at submission (on train soundscapes), i noticed large flocks of grhowl. From there post-processing based on false_positive / true_positive ratio was born. grhowl was the sole, giant, woo-hoo-ing offender. Any sufficiently noisy record was grhowl. Rest of the birds were significantly more behaved, and even if i slightly reduced some birds in the end, it made almost no dfiference to just removing grhowl. That improved score by ~0.015.\n\nNext were experiments with distance and time. First, most simple attempt was finding minimum distance from short train audios to recording locations and predicting zero for any bird above some threshold. That kinda worked, but improvement was minimal (~0.005)\n\nNext i decided to take distance of n% closest samples, instead of the closest one - for example, for 500 records birds i took low ~2%, and kept the distance of 10th closest sample - to avoid outliers. This worked significantly better, and improved score by another ~0.015.\n\nLinear threshold is not based on something real, though. In real life birds do not hit the edge of the map and just stop there. The next iteration of post-processing used gradually increasing probability from ~7 degrees distance to ~2 degrees distance, and it worked better - another ~0.01.\n\nAt the end, time was added to the same code as third dimension (with appropriate scale factor), and all settings were tuned on train soundscapes. Because amount of data points was not very big, i was forced to look at low ~6% of the samples when determining \"distance\" and increase said \"distance\" by quite a bit. It still improved score by another ~0.005.\n\n**Threshold choosing**:\n\nTrain soundscapes and public test set had significantly different optimal thresholds. That alerted me to the possibility of large threshold swings on the full private set. I estimated the final optimal threshold would probably be between them, but slightly to the train soundscapes side, and decided to play it safe. Two different thresholds were chosen - one to intentionally undershoot by a bit, one to intentionally overshoot by a bit - but still at the points where validation scores were relatively flat and only slightly lower than optimal one.\n\nI expected to lose ~0.005-0.01 from that decision in exchange for relative safety from shakeups, but conservative high threshold actually landed on the exact optimal spot of the private LB data.\n\n**Conclusion**:\n\nThat was a fun month and a fun journey from 0.54 private LB to 0.63 private LB. I did not spend a lot of time on this competiton compared to Rainforest ones, but the result was as good.\n\nFinal version of postprocessing (grhowl is on by default, it might be just my models quirk) - https://www.kaggle.com/fffrrt/location-and-time-postprocessing-birdclef-2021\n\nMy best submission scored 0.63/0.68 with it, and 0.59/0.62 without it (with the same threshold selection process)."
  }
}