{
  "id": 178099,
  "title": "Do you use a CNN over a row waveforms?",
  "url": "/competitions/birdsong-recognition/discussion/178099",
  "author_name": "",
  "post_date": "2020-08-28T15:36:18.077370600Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>There are some ways to do it (description from <a href=\"https://arxiv.org/pdf/1912.10211.pdf):\" target=\"_blank\">https://arxiv.org/pdf/1912.10211.pdf):</a></p>\n<p><strong>DaiNet</strong></p>\n<p>We can just exchange a mel-spectrogram with a convolution layer with a relatively large kernel size (80 in this paper) and hope that this size will be enough for describe a central point. It followed by similar CNN architecture, only 1D instead of 2D:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F001c06c6f096404f61684b1859a84568%2FDaiNet.png?generation=1598627341056991&amp;alt=media\" alt=\"\"><br>\nA describing some of the architectures:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7e411aaa9e127668478c27e3ca860893%2FDaiNet1.png?generation=1598626854288543&amp;alt=media\" alt=\"\"><br>\npaper: <a href=\"https://arxiv.org/pdf/1610.00087.pdf\" target=\"_blank\">https://arxiv.org/pdf/1610.00087.pdf</a></p>\n<p><strong>LeeNet</strong></p>\n<p>This paper introduce some splitting of approaches for audio processing:</p>\n<ul>\n<li><em>Frame-level mel-spectrogram model</em>: mel-spectrogram using.</li>\n<li><em>Frame-level raw waveform model</em>: using raw waveforms such that first output shape is equal mel-spectrogram (similar DaiNet in some degree):<ul>\n<li>stride = hop size</li>\n<li>filter length = window size</li>\n<li>number of filters = number of mel-bands</li></ul></li>\n<li><em>Sample-level raw waveform model</em>: using raw waveforms with small kernel sizes all time.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F2674feb80caa13ad9610f55fe62faa77%2FLeeNet1.png?generation=1598627736320475&amp;alt=media\" alt=\"\"><br>\nOne of the best models (3 - kernel and pooling lengths, 9 - depth):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F9b8fb37bb52a3dea7aef4d5474900249%2FLeeNet2.png?generation=1598627840061086&amp;alt=media\" alt=\"\"><br>\npaper: <a href=\"https://arxiv.org/pdf/1703.01789.pdf\" target=\"_blank\">https://arxiv.org/pdf/1703.01789.pdf</a></p>\n<p><strong>Wavegram-CNN</strong></p>\n<p>CNN over row waveform works not so good cause in each moment they have an access to only 1D segment and ignore any signal's frequencies. So we can turn our networks from 1D to 2D dimensions only with reshaping and hope that corresponding matrices can learn some information about frequencies (like mel-spectogramm) or even something better. That's left branch of the picture below and it's name is <em>Wavegram-CNN</em>. And we can reunite this approach with classic mel-spectrogram one and get the net called <em>Wavegram-Logmel-CNN</em>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fc87e5cbadfe2743a46ac89f7b5d0f392%2FWaveGramm.png?generation=1598628362204150&amp;alt=media\" alt=\"\"></p>\n<p>This net achieve the best performance on a huge AudioSet dataset. In my experiments inside this competition <em>Wavegram-Logmel-CNN</em> has achieved one the best results too.</p>\n<p>paper: <a href=\"https://arxiv.org/pdf/1912.10211.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.10211.pdf</a></p>\n<p>Do you use something similar?</p>",
  "messages": [
    {
      "id": "989182",
      "postDate": "08/28/2020 15:36:18",
      "content": "<p>There are some ways to do it (description from <a href=\"https://arxiv.org/pdf/1912.10211.pdf):\" target=\"_blank\">https://arxiv.org/pdf/1912.10211.pdf):</a></p>\n<p><strong>DaiNet</strong></p>\n<p>We can just exchange a mel-spectrogram with a convolution layer with a relatively large kernel size (80 in this paper) and hope that this size will be enough for describe a central point. It followed by similar CNN architecture, only 1D instead of 2D:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F001c06c6f096404f61684b1859a84568%2FDaiNet.png?generation=1598627341056991&amp;alt=media\" alt=\"\"><br>\nA describing some of the architectures:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7e411aaa9e127668478c27e3ca860893%2FDaiNet1.png?generation=1598626854288543&amp;alt=media\" alt=\"\"><br>\npaper: <a href=\"https://arxiv.org/pdf/1610.00087.pdf\" target=\"_blank\">https://arxiv.org/pdf/1610.00087.pdf</a></p>\n<p><strong>LeeNet</strong></p>\n<p>This paper introduce some splitting of approaches for audio processing:</p>\n<ul>\n<li><em>Frame-level mel-spectrogram model</em>: mel-spectrogram using.</li>\n<li><em>Frame-level raw waveform model</em>: using raw waveforms such that first output shape is equal mel-spectrogram (similar DaiNet in some degree):<ul>\n<li>stride = hop size</li>\n<li>filter length = window size</li>\n<li>number of filters = number of mel-bands</li></ul></li>\n<li><em>Sample-level raw waveform model</em>: using raw waveforms with small kernel sizes all time.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F2674feb80caa13ad9610f55fe62faa77%2FLeeNet1.png?generation=1598627736320475&amp;alt=media\" alt=\"\"><br>\nOne of the best models (3 - kernel and pooling lengths, 9 - depth):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F9b8fb37bb52a3dea7aef4d5474900249%2FLeeNet2.png?generation=1598627840061086&amp;alt=media\" alt=\"\"><br>\npaper: <a href=\"https://arxiv.org/pdf/1703.01789.pdf\" target=\"_blank\">https://arxiv.org/pdf/1703.01789.pdf</a></p>\n<p><strong>Wavegram-CNN</strong></p>\n<p>CNN over row waveform works not so good cause in each moment they have an access to only 1D segment and ignore any signal's frequencies. So we can turn our networks from 1D to 2D dimensions only with reshaping and hope that corresponding matrices can learn some information about frequencies (like mel-spectogramm) or even something better. That's left branch of the picture below and it's name is <em>Wavegram-CNN</em>. And we can reunite this approach with classic mel-spectrogram one and get the net called <em>Wavegram-Logmel-CNN</em>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fc87e5cbadfe2743a46ac89f7b5d0f392%2FWaveGramm.png?generation=1598628362204150&amp;alt=media\" alt=\"\"></p>\n<p>This net achieve the best performance on a huge AudioSet dataset. In my experiments inside this competition <em>Wavegram-Logmel-CNN</em> has achieved one the best results too.</p>\n<p>paper: <a href=\"https://arxiv.org/pdf/1912.10211.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.10211.pdf</a></p>\n<p>Do you use something similar?</p>",
      "rawMarkdown": "There are some ways to do it (description from https://arxiv.org/pdf/1912.10211.pdf):\n\n**DaiNet**\n\nWe can just exchange a mel-spectrogram with a convolution layer with a relatively large kernel size (80 in this paper) and hope that this size will be enough for describe a central point. It followed by similar CNN architecture, only 1D instead of 2D:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F001c06c6f096404f61684b1859a84568%2FDaiNet.png?generation=1598627341056991&alt=media)\nA describing some of the architectures:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7e411aaa9e127668478c27e3ca860893%2FDaiNet1.png?generation=1598626854288543&alt=media)\npaper: https://arxiv.org/pdf/1610.00087.pdf\n\n**LeeNet**\n\nThis paper introduce some splitting of approaches for audio processing:\n- *Frame-level mel-spectrogram model*: mel-spectrogram using.\n- *Frame-level raw waveform model*: using raw waveforms such that first output shape is equal mel-spectrogram (similar DaiNet in some degree):\n  - stride = hop size\n  - filter length = window size\n  - number of filters = number of mel-bands\n- *Sample-level raw waveform model*: using raw waveforms with small kernel sizes all time.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F2674feb80caa13ad9610f55fe62faa77%2FLeeNet1.png?generation=1598627736320475&alt=media)\nOne of the best models (3 - kernel and pooling lengths, 9 - depth):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F9b8fb37bb52a3dea7aef4d5474900249%2FLeeNet2.png?generation=1598627840061086&alt=media)\npaper: https://arxiv.org/pdf/1703.01789.pdf\n\n**Wavegram-CNN**\n\nCNN over row waveform works not so good cause in each moment they have an access to only 1D segment and ignore any signal's frequencies. So we can turn our networks from 1D to 2D dimensions only with reshaping and hope that corresponding matrices can learn some information about frequencies (like mel-spectogramm) or even something better. That's left branch of the picture below and it's name is *Wavegram-CNN*. And we can reunite this approach with classic mel-spectrogram one and get the net called *Wavegram-Logmel-CNN*:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fc87e5cbadfe2743a46ac89f7b5d0f392%2FWaveGramm.png?generation=1598628362204150&alt=media)\n\nThis net achieve the best performance on a huge AudioSet dataset. In my experiments inside this competition *Wavegram-Logmel-CNN* has achieved one the best results too.\n\npaper: https://arxiv.org/pdf/1912.10211.pdf\n\nDo you use something similar?",
      "votes": null
    },
    {
      "id": "992154",
      "postDate": "08/31/2020 01:50:18",
      "content": "<p>I tried using 1D CNN on raw audio but the training process took too long and the result didn't meet my expectation.</p>",
      "rawMarkdown": "I tried using 1D CNN on raw audio but the training process took too long and the result didn't meet my expectation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 992154,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/31/2020 01:50:18",
      "content": "<p>I tried using 1D CNN on raw audio but the training process took too long and the result didn't meet my expectation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "989182": "There are some ways to do it (description from https://arxiv.org/pdf/1912.10211.pdf):\n\n**DaiNet**\n\nWe can just exchange a mel-spectrogram with a convolution layer with a relatively large kernel size (80 in this paper) and hope that this size will be enough for describe a central point. It followed by similar CNN architecture, only 1D instead of 2D:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F001c06c6f096404f61684b1859a84568%2FDaiNet.png?generation=1598627341056991&alt=media)\nA describing some of the architectures:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F7e411aaa9e127668478c27e3ca860893%2FDaiNet1.png?generation=1598626854288543&alt=media)\npaper: https://arxiv.org/pdf/1610.00087.pdf\n\n**LeeNet**\n\nThis paper introduce some splitting of approaches for audio processing:\n- *Frame-level mel-spectrogram model*: mel-spectrogram using.\n- *Frame-level raw waveform model*: using raw waveforms such that first output shape is equal mel-spectrogram (similar DaiNet in some degree):\n  - stride = hop size\n  - filter length = window size\n  - number of filters = number of mel-bands\n- *Sample-level raw waveform model*: using raw waveforms with small kernel sizes all time.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F2674feb80caa13ad9610f55fe62faa77%2FLeeNet1.png?generation=1598627736320475&alt=media)\nOne of the best models (3 - kernel and pooling lengths, 9 - depth):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F9b8fb37bb52a3dea7aef4d5474900249%2FLeeNet2.png?generation=1598627840061086&alt=media)\npaper: https://arxiv.org/pdf/1703.01789.pdf\n\n**Wavegram-CNN**\n\nCNN over row waveform works not so good cause in each moment they have an access to only 1D segment and ignore any signal's frequencies. So we can turn our networks from 1D to 2D dimensions only with reshaping and hope that corresponding matrices can learn some information about frequencies (like mel-spectogramm) or even something better. That's left branch of the picture below and it's name is *Wavegram-CNN*. And we can reunite this approach with classic mel-spectrogram one and get the net called *Wavegram-Logmel-CNN*:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fc87e5cbadfe2743a46ac89f7b5d0f392%2FWaveGramm.png?generation=1598628362204150&alt=media)\n\nThis net achieve the best performance on a huge AudioSet dataset. In my experiments inside this competition *Wavegram-Logmel-CNN* has achieved one the best results too.\n\npaper: https://arxiv.org/pdf/1912.10211.pdf\n\nDo you use something similar?",
    "992154": "I tried using 1D CNN on raw audio but the training process took too long and the result didn't meet my expectation."
  },
  "source": "meta"
}