{
  "id": 56720,
  "title": "Deep Learning approach",
  "url": "/competitions/freesound-audio-tagging/discussion/56720",
  "author_name": "",
  "post_date": "2018-05-14T02:48:13.045244100Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, i'm a deep learning enthusiast, and i find audio processing very interesting, so i was thinking about using deep learning on this competition, at first i was going to try RNN (breaking the audio in slices of time), but seems that people are using more CNN (after transforming audio to an image), my question is which is the better option? should i try both?</p>",
  "messages": [
    {
      "id": "328341",
      "postDate": "05/14/2018 02:48:13",
      "content": "<p>Hi, i'm a deep learning enthusiast, and i find audio processing very interesting, so i was thinking about using deep learning on this competition, at first i was going to try RNN (breaking the audio in slices of time), but seems that people are using more CNN (after transforming audio to an image), my question is which is the better option? should i try both?</p>",
      "rawMarkdown": "Hi, i'm a deep learning enthusiast, and i find audio processing very interesting, so i was thinking about using deep learning on this competition, at first i was going to try RNN (breaking the audio in slices of time), but seems that people are using more CNN (after transforming audio to an image), my question is which is the better option? should i try both?",
      "votes": null
    },
    {
      "id": "328352",
      "postDate": "05/14/2018 03:48:04",
      "content": "<p><a href=\"https://arxiv.org/abs/1711.07128\">Hello Edge: Keyword Spotting on Microcontrollers, arXiv:1711.07128 [cs.SD]</a></p>\n\n<p>This paper would be good answer to your question, covers many different models with coherent evaluation.\nIt seems to be anything works, and their best is DS-CNN (Depthwise Separable Convolutional NN).</p>\n\n<p>I'm also using it, though training it is much slower than training AlexNet.\nBoth AlexNet and DS-CNN shows good results.\nCurrently my ensemble of many CNN has LB score 0.916.</p>",
      "rawMarkdown": "[Hello Edge: Keyword Spotting on Microcontrollers, arXiv:1711.07128 [cs.SD]][1]\n\nThis paper would be good answer to your question, covers many different models with coherent evaluation.\nIt seems to be anything works, and their best is DS-CNN (Depthwise Separable Convolutional NN).\n\nI'm also using it, though training it is much slower than training AlexNet.\nBoth AlexNet and DS-CNN shows good results.\nCurrently my ensemble of many CNN has LB score 0.916.\n\n  [1]: https://arxiv.org/abs/1711.07128",
      "votes": null
    },
    {
      "id": "328625",
      "postDate": "05/14/2018 18:28:43",
      "content": "<p>thanks <a href=\"/daisukelab\">@daisukelab</a>, my first idea on this competition was to practice with RNNs, but i will probably end trying both.</p>",
      "rawMarkdown": "thanks @daisukelab, my first idea on this competition was to practice with RNNs, but i will probably end trying both.",
      "votes": null
    },
    {
      "id": "329681",
      "postDate": "05/17/2018 01:10:54",
      "content": "<p>For anyone else interested, i found out that Baidu's deep speech 2 uses a ensemble of CNN and RNN (<a href=\"https://image.slidesharecdn.com/dlsl2017d3l6end-to-endspeechrecognitionwithrecurrentneuralnetworks-170127183704/95/endtoend-speech-recognition-with-recurrent-neural-networks-d3l6-deep-learning-for-speech-and-language-upc-2017-18-638.jpg?cb=1485542507\">as can be seen here</a>), so seems that combining both can be a good approach as well, deep speech 2 is used for another purpose ( Speech-To-Text), but it maybe be worth a try.</p>",
      "rawMarkdown": "For anyone else interested, i found out that Baidu's deep speech 2 uses a ensemble of CNN and RNN ([as can be seen here][1]), so seems that combining both can be a good approach as well, deep speech 2 is used for another purpose ( Speech-To-Text), but it maybe be worth a try.\n\n\n  [1]: https://image.slidesharecdn.com/dlsl2017d3l6end-to-endspeechrecognitionwithrecurrentneuralnetworks-170127183704/95/endtoend-speech-recognition-with-recurrent-neural-networks-d3l6-deep-learning-for-speech-and-language-upc-2017-18-638.jpg?cb=1485542507",
      "votes": null
    },
    {
      "id": "336137",
      "postDate": "05/31/2018 05:59:16",
      "content": "<p>Are you training the model on GPU?</p>",
      "rawMarkdown": "Are you training the model on GPU?",
      "votes": null
    },
    {
      "id": "336170",
      "postDate": "05/31/2018 07:07:57",
      "content": "<p>Yes of course. 500 steps training of DS-CNN took 30 hours for example, ...painful!</p>",
      "rawMarkdown": "Yes of course. 500 steps training of DS-CNN took 30 hours for example, ...painful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 328352,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "05/14/2018 03:48:04",
      "content": "<p><a href=\"https://arxiv.org/abs/1711.07128\">Hello Edge: Keyword Spotting on Microcontrollers, arXiv:1711.07128 [cs.SD]</a></p>\n\n<p>This paper would be good answer to your question, covers many different models with coherent evaluation.\nIt seems to be anything works, and their best is DS-CNN (Depthwise Separable Convolutional NN).</p>\n\n<p>I'm also using it, though training it is much slower than training AlexNet.\nBoth AlexNet and DS-CNN shows good results.\nCurrently my ensemble of many CNN has LB score 0.916.</p>",
      "votes": null,
      "replies": [
        {
          "id": 336137,
          "author_name": "sunjac",
          "author_url": "",
          "post_date": "05/31/2018 05:59:16",
          "content": "<p>Are you training the model on GPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 336170,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/31/2018 07:07:57",
          "content": "<p>Yes of course. 500 steps training of DS-CNN took 30 hours for example, ...painful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 328625,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "05/14/2018 18:28:43",
      "content": "<p>thanks <a href=\"/daisukelab\">@daisukelab</a>, my first idea on this competition was to practice with RNNs, but i will probably end trying both.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 329681,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "05/17/2018 01:10:54",
      "content": "<p>For anyone else interested, i found out that Baidu's deep speech 2 uses a ensemble of CNN and RNN (<a href=\"https://image.slidesharecdn.com/dlsl2017d3l6end-to-endspeechrecognitionwithrecurrentneuralnetworks-170127183704/95/endtoend-speech-recognition-with-recurrent-neural-networks-d3l6-deep-learning-for-speech-and-language-upc-2017-18-638.jpg?cb=1485542507\">as can be seen here</a>), so seems that combining both can be a good approach as well, deep speech 2 is used for another purpose ( Speech-To-Text), but it maybe be worth a try.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "328341": "Hi, i'm a deep learning enthusiast, and i find audio processing very interesting, so i was thinking about using deep learning on this competition, at first i was going to try RNN (breaking the audio in slices of time), but seems that people are using more CNN (after transforming audio to an image), my question is which is the better option? should i try both?",
    "328352": "[Hello Edge: Keyword Spotting on Microcontrollers, arXiv:1711.07128 [cs.SD]][1]\n\nThis paper would be good answer to your question, covers many different models with coherent evaluation.\nIt seems to be anything works, and their best is DS-CNN (Depthwise Separable Convolutional NN).\n\nI'm also using it, though training it is much slower than training AlexNet.\nBoth AlexNet and DS-CNN shows good results.\nCurrently my ensemble of many CNN has LB score 0.916.\n\n  [1]: https://arxiv.org/abs/1711.07128",
    "328625": "thanks @daisukelab, my first idea on this competition was to practice with RNNs, but i will probably end trying both.",
    "329681": "For anyone else interested, i found out that Baidu's deep speech 2 uses a ensemble of CNN and RNN ([as can be seen here][1]), so seems that combining both can be a good approach as well, deep speech 2 is used for another purpose ( Speech-To-Text), but it maybe be worth a try.\n\n\n  [1]: https://image.slidesharecdn.com/dlsl2017d3l6end-to-endspeechrecognitionwithrecurrentneuralnetworks-170127183704/95/endtoend-speech-recognition-with-recurrent-neural-networks-d3l6-deep-learning-for-speech-and-language-upc-2017-18-638.jpg?cb=1485542507",
    "336137": "Are you training the model on GPU?",
    "336170": "Yes of course. 500 steps training of DS-CNN took 30 hours for example, ...painful!"
  },
  "source": "meta"
}