{
  "id": 62634,
  "title": "4th solution",
  "url": "/competitions/freesound-audio-tagging/writeups/nudt-4th-solution",
  "author_name": "",
  "post_date": "2019-01-07T06:12:58.350Z",
  "votes": 41,
  "comment_count": 2,
  "views": 0,
  "content": "<p>For the solution, we employed both deep learning methods and statistic features-based shallow architecture learners. For single model, only deep learning approaches are investigated, and different deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). For log Mel and MFCC, the delta and delta-delta information are also used to formulate three-channels features. Inception, ResNet, ResNeXt, Dual Path Networks (DPN) are selected as the neural network architectures, while Mixup is used for the data augmentation.</p>\n\n<p>Using ResNeXt, our best single convolutional neural network architecture provides a mAP@3 of 0.967 on the public Kaggle leaderboard,  0.942 on private. Moreover, to improve the accuracy further, we also propose a meta learning-based ensemble method. By employing the diversities between different architectures, the meta learning-based model can provide higher prediction accuracy and robustness with the comparison to the single model. Using the proposed meta-learning method, our solution achieves a mAP@3 of 0.977 on the public Kaggle leaderboard, 0.951 on the private LB.</p>\n\n<p>PS: You can find our code given in\n<a href=\"https://github.com/Cocoxili/DCASE2018Task2\">https://github.com/Cocoxili/DCASE2018Task2</a></p>\n\n<p>Please cite this work in your pulications if it helps your research.</p>\n\n<p>@article{xu2018general,\n  title={General audio tagging with ensembling convolutional neural network and statistical features},\n  author={Xu, Kele and Zhu, Boqing and Kong, Qiuqiang and Mi, Haibo and Ding, Bo and Wang, Dezhi and Wang, Huaimin},\n  journal={arXiv preprint arXiv:1810.12832},\n  year={2018}\n}</p>",
  "messages": [
    {
      "id": "366278",
      "postDate": "08/04/2018 16:12:23",
      "content": "<p>For the solution, we employed both deep learning methods and statistic features-based shallow architecture learners. For single model, only deep learning approaches are investigated, and different deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). For log Mel and MFCC, the delta and delta-delta information are also used to formulate three-channels features. Inception, ResNet, ResNeXt, Dual Path Networks (DPN) are selected as the neural network architectures, while Mixup is used for the data augmentation.</p>\n\n<p>Using ResNeXt, our best single convolutional neural network architecture provides a mAP@3 of 0.967 on the public Kaggle leaderboard,  0.942 on private. Moreover, to improve the accuracy further, we also propose a meta learning-based ensemble method. By employing the diversities between different architectures, the meta learning-based model can provide higher prediction accuracy and robustness with the comparison to the single model. Using the proposed meta-learning method, our solution achieves a mAP@3 of 0.977 on the public Kaggle leaderboard, 0.951 on the private LB.</p>\n\n<p>PS: You can find our code given in\n<a href=\"https://github.com/Cocoxili/DCASE2018Task2\">https://github.com/Cocoxili/DCASE2018Task2</a></p>\n\n<p>Please cite this work in your pulications if it helps your research.</p>\n\n<p>@article{xu2018general,\n  title={General audio tagging with ensembling convolutional neural network and statistical features},\n  author={Xu, Kele and Zhu, Boqing and Kong, Qiuqiang and Mi, Haibo and Ding, Bo and Wang, Dezhi and Wang, Huaimin},\n  journal={arXiv preprint arXiv:1810.12832},\n  year={2018}\n}</p>",
      "rawMarkdown": "For the solution, we employed both deep learning methods and statistic features-based shallow architecture learners. For single model, only deep learning approaches are investigated, and different deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). For log Mel and MFCC, the delta and delta-delta information are also used to formulate three-channels features. Inception, ResNet, ResNeXt, Dual Path Networks (DPN) are selected as the neural network architectures, while Mixup is used for the data augmentation.\n\nUsing ResNeXt, our best single convolutional neural network architecture provides a mAP@3 of 0.967 on the public Kaggle leaderboard,  0.942 on private. Moreover, to improve the accuracy further, we also propose a meta learning-based ensemble method. By employing the diversities between different architectures, the meta learning-based model can provide higher prediction accuracy and robustness with the comparison to the single model. Using the proposed meta-learning method, our solution achieves a mAP@3 of 0.977 on the public Kaggle leaderboard, 0.951 on the private LB.\n\nPS: You can find our code given in\nhttps://github.com/Cocoxili/DCASE2018Task2\n\nPlease cite this work in your pulications if it helps your research.\n\n@article{xu2018general,\n  title={General audio tagging with ensembling convolutional neural network and statistical features},\n  author={Xu, Kele and Zhu, Boqing and Kong, Qiuqiang and Mi, Haibo and Ding, Bo and Wang, Dezhi and Wang, Huaimin},\n  journal={arXiv preprint arXiv:1810.12832},\n  year={2018}\n}",
      "votes": null
    },
    {
      "id": "366407",
      "postDate": "08/05/2018 03:37:22",
      "content": "<p>That's interesting... My feature engineering process seems similar to yours. And so is the meta-learning scheme. The models are different though. Congratulations for the wonderful performance. </p>",
      "rawMarkdown": "That's interesting... My feature engineering process seems similar to yours. And so is the meta-learning scheme. The models are different though. Congratulations for the wonderful performance.",
      "votes": null
    },
    {
      "id": "367166",
      "postDate": "08/07/2018 08:15:56",
      "content": "<p>Congratulations! I can't wait to study your code.</p>",
      "rawMarkdown": "Congratulations! I can't wait to study your code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 366407,
      "author_name": "gyat2017",
      "author_url": "",
      "post_date": "08/05/2018 03:37:22",
      "content": "<p>That's interesting... My feature engineering process seems similar to yours. And so is the meta-learning scheme. The models are different though. Congratulations for the wonderful performance. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 367166,
      "author_name": "scutwzm",
      "author_url": "",
      "post_date": "08/07/2018 08:15:56",
      "content": "<p>Congratulations! I can't wait to study your code.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "366278": "For the solution, we employed both deep learning methods and statistic features-based shallow architecture learners. For single model, only deep learning approaches are investigated, and different deep neural network architectures are tested with different kinds of input, which ranges from the raw-signal, log-scaled Mel-spectrograms (log Mel) to Mel Frequency Cepstral Coefficients (MFCC). For log Mel and MFCC, the delta and delta-delta information are also used to formulate three-channels features. Inception, ResNet, ResNeXt, Dual Path Networks (DPN) are selected as the neural network architectures, while Mixup is used for the data augmentation.\n\nUsing ResNeXt, our best single convolutional neural network architecture provides a mAP@3 of 0.967 on the public Kaggle leaderboard,  0.942 on private. Moreover, to improve the accuracy further, we also propose a meta learning-based ensemble method. By employing the diversities between different architectures, the meta learning-based model can provide higher prediction accuracy and robustness with the comparison to the single model. Using the proposed meta-learning method, our solution achieves a mAP@3 of 0.977 on the public Kaggle leaderboard, 0.951 on the private LB.\n\nPS: You can find our code given in\nhttps://github.com/Cocoxili/DCASE2018Task2\n\nPlease cite this work in your pulications if it helps your research.\n\n@article{xu2018general,\n  title={General audio tagging with ensembling convolutional neural network and statistical features},\n  author={Xu, Kele and Zhu, Boqing and Kong, Qiuqiang and Mi, Haibo and Ding, Bo and Wang, Dezhi and Wang, Huaimin},\n  journal={arXiv preprint arXiv:1810.12832},\n  year={2018}\n}",
    "366407": "That's interesting... My feature engineering process seems similar to yours. And so is the meta-learning scheme. The models are different though. Congratulations for the wonderful performance.",
    "367166": "Congratulations! I can't wait to study your code."
  },
  "source": "meta"
}