{
  "id": 183335,
  "title": "My first audio competition but not the last",
  "url": "/competitions/birdsong-recognition/discussion/183335",
  "author_name": "",
  "post_date": "2020-09-16T10:02:44.916376300Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Congratz to the winners, great work!<br>\nThanks to the creators of public kernels and for all discussions, always of high value, teamwork! 🙏🙂<br>\nAnd also thanks to the competition host, perfect challenge 👌</p>\n<hr>\n<p>My first journey with audio data and spectrograms, and what a journey.<br>\nThis competition was a path of learning, new areas all the time, one must be happy with the bronze medal.</p>\n<p><strong>Models/inference:</strong><br>\nStacking resnest50_fast_1s1x64d with B3ns worked best, finally stacked 7 models all with different training method for every fold and with different amount av data, to fit the time limit but to also have a good mix for stacking.<br>\nAlso ensembled two of the best stacked models just to be safe. Also picked best models from both bce_val_loss and F1_samples.</p>\n<p><strong>Training:</strong><br>\nStarted with a basic augmentation Flip H/V/HV and random lines and erasing, worked well against overfitting/underfitting, it's always interesting to see what works even if one can't explain it in detail.<br>\nUsed Mixup, RandomBatchCutout and DropBlock as a base through all testing.<br>\nTested many of the audio and spectrogram techniques but had the best result with the basic augmentation that I started with.<br>\nI think a stayed too long in the Audio/Image/Spec Augmentation test/play-zone, could have used that time for other and better competition tasks, but the knowledge from that reading and testing,I will certainly have use of in the future.</p>\n<p>Using optimizer and scheduler AdamP, CosineAnnealingLR (8 epoch cycles within LR 8e-4 and 1e-4) had the best train/valid curve and result with all models.</p>\n<p>Last week I started using different losses like BCEWithLogitsLoss(reduction='sum') and LSEPloss instead of vanilla  BCEWithLogitsLoss, and augmentation techniques from prev. competitions like below, maybe they could have worked better with more time and testing.<br>\nAlso started using mixed precision way too late.</p>\n<p>transforms_train = Compose([<br>\n        OneOf([<br>\n            PadToSize(size, mode='wrap'),<br>\n            PadToSize(size, mode='constant')<br>\n        ], p=[0.5, 0.5]),<br>\n       # RandomCrop(size),<br>\n        UseWithProb(<br>\n            RandomResizedCrop(scale=(0.8, 1.0), ratio=(1.7, 2.3)),<br>\n            prob=0.33<br>\n        ),<br>\n        UseWithProb(SpecAugment(num_mask=2,<br>\n                                freq_masking=0.15,<br>\n                                time_masking=0.20), 0.5),<br>\n                ])</p>\n<p><strong>Other interesting notes…:</strong><br>\nFirst trained a model with 1/5 split and then finetune it with 1/9 split, smaller validset within the prev one, worked well again, also tried it in the last competition, time saving.</p>\n<p>Also tested upsampling but didn't work for me, never has in any competition, maybe better with right loss/weights instead.</p>\n<p>F1 optimizer:<br>\nTested a solution to scan and optimize the local threshold, it created thresholds like below, perfect on specific OOF but not to the final submission. Didn't have time finish the code, maybe next time.<br>\narray([0.995, 0.995, 0.955, 0.995, 0.98 , 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.785, 0.995,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.935, 0.995,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.965, 0.995,<br>\n       0.69 , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.99 , 0.   , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.995, 0.995, 0.945, 0.995, 0.995, 0.995, 0.995, 0.98 ,<br>\n       0.995, 0.985, 0.995, 0.995, 0.995, 0.99 , 0.995, 0.995, 0.955,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.97 , 0.995, 0.99 , 0.895, 0.995, 0.995, 0.975, 0.995,<br>\n       0.995, 0.995, 0.995, 0.985, 0.995, 0.995, 0.92 , 0.995, 0.995,<br>\n       0.995, 0.975, 0.995, 0.795, 0.995, 0.995, 0.995, 0.995, 0.995,…………..</p>\n<p><strong>Conclusion</strong>:<br>\nThink I got the best out of the setup that I used, with more time, more various of models, data, spectrogram setups should have been tested.<br>\nNice to see in the end though , after a lot of testing, a solid F1 sample scores, 0.569-0.574 Public, 0.57-0.62 CV, 0.594-0.604 Private.<br>\nNow I have at least a good base of knowledge when the next Audio challenge rises on the horizon 👍</p>\n<p><strong>Until next time</strong><br>\n ….did one think of all the call variations that exists..</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F88c3edb5e2925567c9e88eead3b95286%2F5e4ff5743fe2ff86bb0bbca80b93b4cd.jpg?generation=1600248861921253&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1012812",
      "postDate": "09/16/2020 10:02:44",
      "content": "<p>Congratz to the winners, great work!<br>\nThanks to the creators of public kernels and for all discussions, always of high value, teamwork! 🙏🙂<br>\nAnd also thanks to the competition host, perfect challenge 👌</p>\n<hr>\n<p>My first journey with audio data and spectrograms, and what a journey.<br>\nThis competition was a path of learning, new areas all the time, one must be happy with the bronze medal.</p>\n<p><strong>Models/inference:</strong><br>\nStacking resnest50_fast_1s1x64d with B3ns worked best, finally stacked 7 models all with different training method for every fold and with different amount av data, to fit the time limit but to also have a good mix for stacking.<br>\nAlso ensembled two of the best stacked models just to be safe. Also picked best models from both bce_val_loss and F1_samples.</p>\n<p><strong>Training:</strong><br>\nStarted with a basic augmentation Flip H/V/HV and random lines and erasing, worked well against overfitting/underfitting, it's always interesting to see what works even if one can't explain it in detail.<br>\nUsed Mixup, RandomBatchCutout and DropBlock as a base through all testing.<br>\nTested many of the audio and spectrogram techniques but had the best result with the basic augmentation that I started with.<br>\nI think a stayed too long in the Audio/Image/Spec Augmentation test/play-zone, could have used that time for other and better competition tasks, but the knowledge from that reading and testing,I will certainly have use of in the future.</p>\n<p>Using optimizer and scheduler AdamP, CosineAnnealingLR (8 epoch cycles within LR 8e-4 and 1e-4) had the best train/valid curve and result with all models.</p>\n<p>Last week I started using different losses like BCEWithLogitsLoss(reduction='sum') and LSEPloss instead of vanilla  BCEWithLogitsLoss, and augmentation techniques from prev. competitions like below, maybe they could have worked better with more time and testing.<br>\nAlso started using mixed precision way too late.</p>\n<p>transforms_train = Compose([<br>\n        OneOf([<br>\n            PadToSize(size, mode='wrap'),<br>\n            PadToSize(size, mode='constant')<br>\n        ], p=[0.5, 0.5]),<br>\n       # RandomCrop(size),<br>\n        UseWithProb(<br>\n            RandomResizedCrop(scale=(0.8, 1.0), ratio=(1.7, 2.3)),<br>\n            prob=0.33<br>\n        ),<br>\n        UseWithProb(SpecAugment(num_mask=2,<br>\n                                freq_masking=0.15,<br>\n                                time_masking=0.20), 0.5),<br>\n                ])</p>\n<p><strong>Other interesting notes…:</strong><br>\nFirst trained a model with 1/5 split and then finetune it with 1/9 split, smaller validset within the prev one, worked well again, also tried it in the last competition, time saving.</p>\n<p>Also tested upsampling but didn't work for me, never has in any competition, maybe better with right loss/weights instead.</p>\n<p>F1 optimizer:<br>\nTested a solution to scan and optimize the local threshold, it created thresholds like below, perfect on specific OOF but not to the final submission. Didn't have time finish the code, maybe next time.<br>\narray([0.995, 0.995, 0.955, 0.995, 0.98 , 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.785, 0.995,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.935, 0.995,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.965, 0.995,<br>\n       0.69 , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.99 , 0.   , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.995, 0.995, 0.945, 0.995, 0.995, 0.995, 0.995, 0.98 ,<br>\n       0.995, 0.985, 0.995, 0.995, 0.995, 0.99 , 0.995, 0.995, 0.955,<br>\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,<br>\n       0.995, 0.97 , 0.995, 0.99 , 0.895, 0.995, 0.995, 0.975, 0.995,<br>\n       0.995, 0.995, 0.995, 0.985, 0.995, 0.995, 0.92 , 0.995, 0.995,<br>\n       0.995, 0.975, 0.995, 0.795, 0.995, 0.995, 0.995, 0.995, 0.995,…………..</p>\n<p><strong>Conclusion</strong>:<br>\nThink I got the best out of the setup that I used, with more time, more various of models, data, spectrogram setups should have been tested.<br>\nNice to see in the end though , after a lot of testing, a solid F1 sample scores, 0.569-0.574 Public, 0.57-0.62 CV, 0.594-0.604 Private.<br>\nNow I have at least a good base of knowledge when the next Audio challenge rises on the horizon 👍</p>\n<p><strong>Until next time</strong><br>\n ….did one think of all the call variations that exists..</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F88c3edb5e2925567c9e88eead3b95286%2F5e4ff5743fe2ff86bb0bbca80b93b4cd.jpg?generation=1600248861921253&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratz to the winners, great work!\nThanks to the creators of public kernels and for all discussions, always of high value, teamwork! 🙏🙂\nAnd also thanks to the competition host, perfect challenge 👌\n\n------------\n\nMy first journey with audio data and spectrograms, and what a journey.\nThis competition was a path of learning, new areas all the time, one must be happy with the bronze medal.\n\n**Models/inference:**\nStacking resnest50_fast_1s1x64d with B3ns worked best, finally stacked 7 models all with different training method for every fold and with different amount av data, to fit the time limit but to also have a good mix for stacking.\nAlso ensembled two of the best stacked models just to be safe. Also picked best models from both bce_val_loss and F1_samples.\n\n**Training:**\nStarted with a basic augmentation Flip H/V/HV and random lines and erasing, worked well against overfitting/underfitting, it's always interesting to see what works even if one can't explain it in detail.\nUsed Mixup, RandomBatchCutout and DropBlock as a base through all testing.\nTested many of the audio and spectrogram techniques but had the best result with the basic augmentation that I started with.\nI think a stayed too long in the Audio/Image/Spec Augmentation test/play-zone, could have used that time for other and better competition tasks, but the knowledge from that reading and testing,I will certainly have use of in the future.\n\nUsing optimizer and scheduler AdamP, CosineAnnealingLR (8 epoch cycles within LR 8e-4 and 1e-4) had the best train/valid curve and result with all models.\n\nLast week I started using different losses like BCEWithLogitsLoss(reduction='sum') and LSEPloss instead of vanilla  BCEWithLogitsLoss, and augmentation techniques from prev. competitions like below, maybe they could have worked better with more time and testing.\nAlso started using mixed precision way too late.\n\ntransforms_train = Compose([\n        OneOf([\n            PadToSize(size, mode='wrap'),\n            PadToSize(size, mode='constant')\n        ], p=[0.5, 0.5]),\n       # RandomCrop(size),\n        UseWithProb(\n            RandomResizedCrop(scale=(0.8, 1.0), ratio=(1.7, 2.3)),\n            prob=0.33\n        ),\n        UseWithProb(SpecAugment(num_mask=2,\n                                freq_masking=0.15,\n                                time_masking=0.20), 0.5),\n                ])\n\n\n**Other interesting notes...:**\nFirst trained a model with 1/5 split and then finetune it with 1/9 split, smaller validset within the prev one, worked well again, also tried it in the last competition, time saving.\n\nAlso tested upsampling but didn't work for me, never has in any competition, maybe better with right loss/weights instead.\n\nF1 optimizer:\nTested a solution to scan and optimize the local threshold, it created thresholds like below, perfect on specific OOF but not to the final submission. Didn't have time finish the code, maybe next time.\narray([0.995, 0.995, 0.955, 0.995, 0.98 , 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.785, 0.995,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.935, 0.995,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.965, 0.995,\n       0.69 , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.99 , 0.   , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.995, 0.995, 0.945, 0.995, 0.995, 0.995, 0.995, 0.98 ,\n       0.995, 0.985, 0.995, 0.995, 0.995, 0.99 , 0.995, 0.995, 0.955,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.97 , 0.995, 0.99 , 0.895, 0.995, 0.995, 0.975, 0.995,\n       0.995, 0.995, 0.995, 0.985, 0.995, 0.995, 0.92 , 0.995, 0.995,\n       0.995, 0.975, 0.995, 0.795, 0.995, 0.995, 0.995, 0.995, 0.995,..............\n\n**Conclusion**:\nThink I got the best out of the setup that I used, with more time, more various of models, data, spectrogram setups should have been tested.\nNice to see in the end though , after a lot of testing, a solid F1 sample scores, 0.569-0.574 Public, 0.57-0.62 CV, 0.594-0.604 Private.\nNow I have at least a good base of knowledge when the next Audio challenge rises on the horizon 👍\n\n\n**Until next time**\n ....did one think of all the call variations that exists..\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F88c3edb5e2925567c9e88eead3b95286%2F5e4ff5743fe2ff86bb0bbca80b93b4cd.jpg?generation=1600248861921253&alt=media)",
      "votes": null
    },
    {
      "id": "1016120",
      "postDate": "09/18/2020 17:17:14",
      "content": "<p>Well done everybody!</p>",
      "rawMarkdown": "Well done everybody!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1016120,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "09/18/2020 17:17:14",
      "content": "<p>Well done everybody!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1012812": "Congratz to the winners, great work!\nThanks to the creators of public kernels and for all discussions, always of high value, teamwork! 🙏🙂\nAnd also thanks to the competition host, perfect challenge 👌\n\n------------\n\nMy first journey with audio data and spectrograms, and what a journey.\nThis competition was a path of learning, new areas all the time, one must be happy with the bronze medal.\n\n**Models/inference:**\nStacking resnest50_fast_1s1x64d with B3ns worked best, finally stacked 7 models all with different training method for every fold and with different amount av data, to fit the time limit but to also have a good mix for stacking.\nAlso ensembled two of the best stacked models just to be safe. Also picked best models from both bce_val_loss and F1_samples.\n\n**Training:**\nStarted with a basic augmentation Flip H/V/HV and random lines and erasing, worked well against overfitting/underfitting, it's always interesting to see what works even if one can't explain it in detail.\nUsed Mixup, RandomBatchCutout and DropBlock as a base through all testing.\nTested many of the audio and spectrogram techniques but had the best result with the basic augmentation that I started with.\nI think a stayed too long in the Audio/Image/Spec Augmentation test/play-zone, could have used that time for other and better competition tasks, but the knowledge from that reading and testing,I will certainly have use of in the future.\n\nUsing optimizer and scheduler AdamP, CosineAnnealingLR (8 epoch cycles within LR 8e-4 and 1e-4) had the best train/valid curve and result with all models.\n\nLast week I started using different losses like BCEWithLogitsLoss(reduction='sum') and LSEPloss instead of vanilla  BCEWithLogitsLoss, and augmentation techniques from prev. competitions like below, maybe they could have worked better with more time and testing.\nAlso started using mixed precision way too late.\n\ntransforms_train = Compose([\n        OneOf([\n            PadToSize(size, mode='wrap'),\n            PadToSize(size, mode='constant')\n        ], p=[0.5, 0.5]),\n       # RandomCrop(size),\n        UseWithProb(\n            RandomResizedCrop(scale=(0.8, 1.0), ratio=(1.7, 2.3)),\n            prob=0.33\n        ),\n        UseWithProb(SpecAugment(num_mask=2,\n                                freq_masking=0.15,\n                                time_masking=0.20), 0.5),\n                ])\n\n\n**Other interesting notes...:**\nFirst trained a model with 1/5 split and then finetune it with 1/9 split, smaller validset within the prev one, worked well again, also tried it in the last competition, time saving.\n\nAlso tested upsampling but didn't work for me, never has in any competition, maybe better with right loss/weights instead.\n\nF1 optimizer:\nTested a solution to scan and optimize the local threshold, it created thresholds like below, perfect on specific OOF but not to the final submission. Didn't have time finish the code, maybe next time.\narray([0.995, 0.995, 0.955, 0.995, 0.98 , 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.785, 0.995,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.935, 0.995,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.965, 0.995,\n       0.69 , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.99 , 0.   , 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.995, 0.995, 0.945, 0.995, 0.995, 0.995, 0.995, 0.98 ,\n       0.995, 0.985, 0.995, 0.995, 0.995, 0.99 , 0.995, 0.995, 0.955,\n       0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995, 0.995,\n       0.995, 0.97 , 0.995, 0.99 , 0.895, 0.995, 0.995, 0.975, 0.995,\n       0.995, 0.995, 0.995, 0.985, 0.995, 0.995, 0.92 , 0.995, 0.995,\n       0.995, 0.975, 0.995, 0.795, 0.995, 0.995, 0.995, 0.995, 0.995,..............\n\n**Conclusion**:\nThink I got the best out of the setup that I used, with more time, more various of models, data, spectrogram setups should have been tested.\nNice to see in the end though , after a lot of testing, a solid F1 sample scores, 0.569-0.574 Public, 0.57-0.62 CV, 0.594-0.604 Private.\nNow I have at least a good base of knowledge when the next Audio challenge rises on the horizon 👍\n\n\n**Until next time**\n ....did one think of all the call variations that exists..\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F88c3edb5e2925567c9e88eead3b95286%2F5e4ff5743fe2ff86bb0bbca80b93b4cd.jpg?generation=1600248861921253&alt=media)",
    "1016120": "Well done everybody!"
  },
  "source": "meta"
}