{
  "id": 318367,
  "title": "Why we are using CNN for audio data?",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/318367",
  "author_name": "",
  "post_date": "2022-04-12T04:17:48.675742400Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have seen CNN being used for audio classification several times and I am not getting the reason behind the same. Well, it performs well on audio but the question still remains.</p>\n<p>Basic intuition behind the invention of CNN was to extract local features from images and it gives quite impressing results on image datasets. But, what was the first intuition to try it on audio data? In other words, why CNN gives great results on audio data as well?</p>\n<p>It will be very helpful if someone can point out where this all began. Any other thoughts/views are welcome!</p>",
  "messages": [
    {
      "id": "1752704",
      "postDate": "04/12/2022 04:17:48",
      "content": "<p>I have seen CNN being used for audio classification several times and I am not getting the reason behind the same. Well, it performs well on audio but the question still remains.</p>\n<p>Basic intuition behind the invention of CNN was to extract local features from images and it gives quite impressing results on image datasets. But, what was the first intuition to try it on audio data? In other words, why CNN gives great results on audio data as well?</p>\n<p>It will be very helpful if someone can point out where this all began. Any other thoughts/views are welcome!</p>",
      "rawMarkdown": "I have seen CNN being used for audio classification several times and I am not getting the reason behind the same. Well, it performs well on audio but the question still remains.\n\nBasic intuition behind the invention of CNN was to extract local features from images and it gives quite impressing results on image datasets. But, what was the first intuition to try it on audio data? In other words, why CNN gives great results on audio data as well?\n\nIt will be very helpful if someone can point out where this all began. Any other thoughts/views are welcome!",
      "votes": null
    },
    {
      "id": "1759325",
      "postDate": "04/18/2022 14:38:38",
      "content": "<p>This article explains why using a mel-spectrogram works so well: <a href=\"https://towardsdatascience.com/audio-deep-learning-made-simple-part-2-why-mel-spectrograms-perform-better-aad889a93505\" target=\"_blank\">Link</a></p>",
      "rawMarkdown": "This article explains why using a mel-spectrogram works so well: [Link](https://towardsdatascience.com/audio-deep-learning-made-simple-part-2-why-mel-spectrograms-perform-better-aad889a93505)",
      "votes": null
    },
    {
      "id": "1762898",
      "postDate": "04/21/2022 04:20:56",
      "content": "<p>This is helpful. Most of the ideas in ML are derived by observing how we, humans, perceives things. It is quite interesting that we can't <strong>exactly</strong> tell what is happening inside DL models but it works well. </p>\n<p>Well, I have already gone through theory of Mel Spectrograms &amp; MFCCs. MFCCs are the most favorite feature for any application of audio data. The intuition behind invention of this feature had also came from how human vocal track generates the sound. The strange thing is that MFCC also gives good results on sounds not generated by humans. </p>\n<p>At the same time CNN gives good results not only while trained on Mel Spectrograms but also when trained on any features (like MFCC) of audio data. So, the question still remains: what was the intuition to apply CNN to audio data!</p>",
      "rawMarkdown": "This is helpful. Most of the ideas in ML are derived by observing how we, humans, perceives things. It is quite interesting that we can't **exactly** tell what is happening inside DL models but it works well. \n\nWell, I have already gone through theory of Mel Spectrograms & MFCCs. MFCCs are the most favorite feature for any application of audio data. The intuition behind invention of this feature had also came from how human vocal track generates the sound. The strange thing is that MFCC also gives good results on sounds not generated by humans. \n\nAt the same time CNN gives good results not only while trained on Mel Spectrograms but also when trained on any features (like MFCC) of audio data. So, the question still remains: what was the intuition to apply CNN to audio data!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1759325,
      "author_name": "pcmiller",
      "author_url": "",
      "post_date": "04/18/2022 14:38:38",
      "content": "<p>This article explains why using a mel-spectrogram works so well: <a href=\"https://towardsdatascience.com/audio-deep-learning-made-simple-part-2-why-mel-spectrograms-perform-better-aad889a93505\" target=\"_blank\">Link</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1762898,
          "author_name": "neelnirajjoshi",
          "author_url": "",
          "post_date": "04/21/2022 04:20:56",
          "content": "<p>This is helpful. Most of the ideas in ML are derived by observing how we, humans, perceives things. It is quite interesting that we can't <strong>exactly</strong> tell what is happening inside DL models but it works well. </p>\n<p>Well, I have already gone through theory of Mel Spectrograms &amp; MFCCs. MFCCs are the most favorite feature for any application of audio data. The intuition behind invention of this feature had also came from how human vocal track generates the sound. The strange thing is that MFCC also gives good results on sounds not generated by humans. </p>\n<p>At the same time CNN gives good results not only while trained on Mel Spectrograms but also when trained on any features (like MFCC) of audio data. So, the question still remains: what was the intuition to apply CNN to audio data!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1752704": "I have seen CNN being used for audio classification several times and I am not getting the reason behind the same. Well, it performs well on audio but the question still remains.\n\nBasic intuition behind the invention of CNN was to extract local features from images and it gives quite impressing results on image datasets. But, what was the first intuition to try it on audio data? In other words, why CNN gives great results on audio data as well?\n\nIt will be very helpful if someone can point out where this all began. Any other thoughts/views are welcome!",
    "1759325": "This article explains why using a mel-spectrogram works so well: [Link](https://towardsdatascience.com/audio-deep-learning-made-simple-part-2-why-mel-spectrograms-perform-better-aad889a93505)",
    "1762898": "This is helpful. Most of the ideas in ML are derived by observing how we, humans, perceives things. It is quite interesting that we can't **exactly** tell what is happening inside DL models but it works well. \n\nWell, I have already gone through theory of Mel Spectrograms & MFCCs. MFCCs are the most favorite feature for any application of audio data. The intuition behind invention of this feature had also came from how human vocal track generates the sound. The strange thing is that MFCC also gives good results on sounds not generated by humans. \n\nAt the same time CNN gives good results not only while trained on Mel Spectrograms but also when trained on any features (like MFCC) of audio data. So, the question still remains: what was the intuition to apply CNN to audio data!"
  },
  "source": "meta"
}