{
  "id": 158943,
  "title": "Previous Audio competitions and solutions",
  "url": "/competitions/birdsong-recognition/discussion/158943",
  "author_name": "",
  "post_date": "2020-06-15T21:01:48.116251100Z",
  "votes": 86,
  "comment_count": 2,
  "views": 0,
  "content": "<h2><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019\">Freesound Audio Tagging 2019</a></h2>\n\n<p>![](<a href=\"https://storage.googleapis.com/kaggle-media/competitions/freesound/task2_freesound_audio_tagging.png\">https://storage.googleapis.com/kaggle-media/competitions/freesound/task2_freesound_audio_tagging.png</a> =250x*)</p>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924\">1st place solution</a></h3>\n\n<ul>\n<li>CNN model with attention, skip connections and auxiliary classifiers</li>\n<li>SpecAugment, Mixup augmentations</li>\n<li>Ensemble with MLP second-level model and geometric mean blending</li>\n</ul>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815\">2nd place solution</a></h3>\n\n<ul>\n<li>Feature engineering: Log-mel (441,64) (time,mels) , global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005), length</li>\n<li>Prepocess: audio clips are first trimmed of leading and trailing silence, random select a 5s clip from audio clip.</li>\n<li>Custom CNN.</li>\n<li>Data Augmentation: Mixup , random select 5s clip + random padding</li>\n</ul>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\">3rd place solution</a></h3>\n\n<ul>\n<li>custom CNNs</li>\n<li>augmentations to be very important for getting good results:  MixUp, audio effects such as reverb, pitch, tempo and overdrive.</li>\n<li>Training: large audio segments for training,  segments from 8 to 12 seconds.</li>\n<li>Inference: No TTA, used full-length audio instead.</li>\n<li>Noisy data:  iterative pseudolabeling</li>\n</ul>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440\">4th place solution</a></h3>\n\n<ul>\n<li>ResNet34 with log-mel / EnvNet-v2 with waveform.</li>\n<li>Multitask learning with noisy labels.</li>\n<li>Semi-supervised learning (SSL) with noisy data.</li>\n<li>Averaging models trained with different time windows.</li>\n<li>augmentations: log-mel , waveform</li>\n</ul>\n\n<hr>\n\n<h2><a href=\"https://www.kaggle.com/c/freesound-audio-tagging\">Freesound General-Purpose Audio Tagging Challenge</a></h2>\n\n<p><strong><a href=\"https://arxiv.org/abs/1810.12832\">Paper</a></strong></p>\n\n<ul>\n<li>Deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). </li>\n<li>Models: Inception, ResNet, ResNeXt (best single), Dual Path Networks (DPN)</li>\n<li>Mixup is used for the data augmentation.</li>\n</ul>\n\n<hr>\n\n<p>hope this helps, please check the original threads from the authors for more information.</p>",
  "messages": [
    {
      "id": "887761",
      "postDate": "06/15/2020 21:01:48",
      "content": "<h2><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019\">Freesound Audio Tagging 2019</a></h2>\n\n<p>![](<a href=\"https://storage.googleapis.com/kaggle-media/competitions/freesound/task2_freesound_audio_tagging.png\">https://storage.googleapis.com/kaggle-media/competitions/freesound/task2_freesound_audio_tagging.png</a> =250x*)</p>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924\">1st place solution</a></h3>\n\n<ul>\n<li>CNN model with attention, skip connections and auxiliary classifiers</li>\n<li>SpecAugment, Mixup augmentations</li>\n<li>Ensemble with MLP second-level model and geometric mean blending</li>\n</ul>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815\">2nd place solution</a></h3>\n\n<ul>\n<li>Feature engineering: Log-mel (441,64) (time,mels) , global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005), length</li>\n<li>Prepocess: audio clips are first trimmed of leading and trailing silence, random select a 5s clip from audio clip.</li>\n<li>Custom CNN.</li>\n<li>Data Augmentation: Mixup , random select 5s clip + random padding</li>\n</ul>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926\">3rd place solution</a></h3>\n\n<ul>\n<li>custom CNNs</li>\n<li>augmentations to be very important for getting good results:  MixUp, audio effects such as reverb, pitch, tempo and overdrive.</li>\n<li>Training: large audio segments for training,  segments from 8 to 12 seconds.</li>\n<li>Inference: No TTA, used full-length audio instead.</li>\n<li>Noisy data:  iterative pseudolabeling</li>\n</ul>\n\n<h3><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440\">4th place solution</a></h3>\n\n<ul>\n<li>ResNet34 with log-mel / EnvNet-v2 with waveform.</li>\n<li>Multitask learning with noisy labels.</li>\n<li>Semi-supervised learning (SSL) with noisy data.</li>\n<li>Averaging models trained with different time windows.</li>\n<li>augmentations: log-mel , waveform</li>\n</ul>\n\n<hr>\n\n<h2><a href=\"https://www.kaggle.com/c/freesound-audio-tagging\">Freesound General-Purpose Audio Tagging Challenge</a></h2>\n\n<p><strong><a href=\"https://arxiv.org/abs/1810.12832\">Paper</a></strong></p>\n\n<ul>\n<li>Deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). </li>\n<li>Models: Inception, ResNet, ResNeXt (best single), Dual Path Networks (DPN)</li>\n<li>Mixup is used for the data augmentation.</li>\n</ul>\n\n<hr>\n\n<p>hope this helps, please check the original threads from the authors for more information.</p>",
      "rawMarkdown": "## [Freesound Audio Tagging 2019](https://www.kaggle.com/c/freesound-audio-tagging-2019)\n\n![](https://storage.googleapis.com/kaggle-media/competitions/freesound/task2_freesound_audio_tagging.png =250x*)\n\n###[1st place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924)\n- CNN model with attention, skip connections and auxiliary classifiers\n- SpecAugment, Mixup augmentations\n- Ensemble with MLP second-level model and geometric mean blending\n\n### [2nd place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815)\n\n- Feature engineering: Log-mel (441,64) (time,mels) , global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005), length\n- Prepocess: audio clips are first trimmed of leading and trailing silence, random select a 5s clip from audio clip.\n- Custom CNN.\n- Data Augmentation: Mixup , random select 5s clip + random padding\n\n### [3rd place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926)\n\n- custom CNNs\n- augmentations to be very important for getting good results:  MixUp, audio effects such as reverb, pitch, tempo and overdrive.\n- Training: large audio segments for training,  segments from 8 to 12 seconds.\n- Inference: No TTA, used full-length audio instead.\n- Noisy data:  iterative pseudolabeling\n\n### [4th place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440)\n- ResNet34 with log-mel / EnvNet-v2 with waveform.\n- Multitask learning with noisy labels.\n- Semi-supervised learning (SSL) with noisy data.\n- Averaging models trained with different time windows.\n- augmentations: log-mel , waveform\n\n----\n\n## [Freesound General-Purpose Audio Tagging Challenge](https://www.kaggle.com/c/freesound-audio-tagging)\n\n**[Paper](https://arxiv.org/abs/1810.12832)**\n\n- Deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). \n- Models: Inception, ResNet, ResNeXt (best single), Dual Path Networks (DPN)\n- Mixup is used for the data augmentation.\n\n---- \n\nhope this helps, please check the original threads from the authors for more information.",
      "votes": null
    },
    {
      "id": "888269",
      "postDate": "06/16/2020 08:42:53",
      "content": "<p>Thank you a lot for sharing!</p>",
      "rawMarkdown": "Thank you a lot for sharing!",
      "votes": null
    },
    {
      "id": "1187334",
      "postDate": "02/05/2021 11:26:11",
      "content": "<p>Good work.<br>\nThanks for sharing this.</p>",
      "rawMarkdown": "Good work.\nThanks for sharing this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1187334,
      "author_name": "tasneemabdulrahim",
      "author_url": "",
      "post_date": "02/05/2021 11:26:11",
      "content": "<p>Good work.<br>\nThanks for sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 888269,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "06/16/2020 08:42:53",
      "content": "<p>Thank you a lot for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "887761": "## [Freesound Audio Tagging 2019](https://www.kaggle.com/c/freesound-audio-tagging-2019)\n\n![](https://storage.googleapis.com/kaggle-media/competitions/freesound/task2_freesound_audio_tagging.png =250x*)\n\n###[1st place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924)\n- CNN model with attention, skip connections and auxiliary classifiers\n- SpecAugment, Mixup augmentations\n- Ensemble with MLP second-level model and geometric mean blending\n\n### [2nd place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97815)\n\n- Feature engineering: Log-mel (441,64) (time,mels) , global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005), length\n- Prepocess: audio clips are first trimmed of leading and trailing silence, random select a 5s clip from audio clip.\n- Custom CNN.\n- Data Augmentation: Mixup , random select 5s clip + random padding\n\n### [3rd place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/97926)\n\n- custom CNNs\n- augmentations to be very important for getting good results:  MixUp, audio effects such as reverb, pitch, tempo and overdrive.\n- Training: large audio segments for training,  segments from 8 to 12 seconds.\n- Inference: No TTA, used full-length audio instead.\n- Noisy data:  iterative pseudolabeling\n\n### [4th place solution](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/96440)\n- ResNet34 with log-mel / EnvNet-v2 with waveform.\n- Multitask learning with noisy labels.\n- Semi-supervised learning (SSL) with noisy data.\n- Averaging models trained with different time windows.\n- augmentations: log-mel , waveform\n\n----\n\n## [Freesound General-Purpose Audio Tagging Challenge](https://www.kaggle.com/c/freesound-audio-tagging)\n\n**[Paper](https://arxiv.org/abs/1810.12832)**\n\n- Deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). \n- Models: Inception, ResNet, ResNeXt (best single), Dual Path Networks (DPN)\n- Mixup is used for the data augmentation.\n\n---- \n\nhope this helps, please check the original threads from the authors for more information.",
    "888269": "Thank you a lot for sharing!",
    "1187334": "Good work.\nThanks for sharing this."
  },
  "source": "meta"
}