{
  "id": 205571,
  "title": "What about using the sound data directly?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/205571",
  "author_name": "",
  "post_date": "2020-12-20T17:38:29.326799300Z",
  "votes": 5,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>first I have to say, that this is my first accoustic data project, so maybe my questions sound stupid to you.<br>\nTo me it seems a bit odd, that we have to produce spectrograms to feed them into a CNN. Ok I know, this is maybe the way a human actually would try to identify a species while hearing to the soud it produces, but don't we really waste a lot of info this way?</p>\n<p>Wouldn't it be practical to directly feed the audio data (probably after converting it to wav) to a CNN?<br>\nI am sure this was already tried a lot of times, because it seems straightforward to me, but the more I wonder why this method has been abandomed by most of the people working on accoustic signals.</p>\n<p>Can anybody shed some light please?</p>\n<p>Thank you in advance<br>\nJürgen</p>",
  "messages": [
    {
      "id": "1120253",
      "postDate": "12/20/2020 17:38:29",
      "content": "<p>Hi,</p>\n<p>first I have to say, that this is my first accoustic data project, so maybe my questions sound stupid to you.<br>\nTo me it seems a bit odd, that we have to produce spectrograms to feed them into a CNN. Ok I know, this is maybe the way a human actually would try to identify a species while hearing to the soud it produces, but don't we really waste a lot of info this way?</p>\n<p>Wouldn't it be practical to directly feed the audio data (probably after converting it to wav) to a CNN?<br>\nI am sure this was already tried a lot of times, because it seems straightforward to me, but the more I wonder why this method has been abandomed by most of the people working on accoustic signals.</p>\n<p>Can anybody shed some light please?</p>\n<p>Thank you in advance<br>\nJürgen</p>",
      "rawMarkdown": "Hi,\n\nfirst I have to say, that this is my first accoustic data project, so maybe my questions sound stupid to you.\nTo me it seems a bit odd, that we have to produce spectrograms to feed them into a CNN. Ok I know, this is maybe the way a human actually would try to identify a species while hearing to the soud it produces, but don't we really waste a lot of info this way?\n\nWouldn't it be practical to directly feed the audio data (probably after converting it to wav) to a CNN?\nI am sure this was already tried a lot of times, because it seems straightforward to me, but the more I wonder why this method has been abandomed by most of the people working on accoustic signals.\n\nCan anybody shed some light please?\n\nThank you in advance\nJürgen",
      "votes": null
    },
    {
      "id": "1120599",
      "postDate": "12/21/2020 00:55:45",
      "content": "<p>I guess, we kind of know that (Mel-)spectrogram are pretty good features. If you know that, it can in a limited data regimen be preferable to use those vs. learning features. You might wonder about just putting raw signal into a LSTM or some kind of transformer, but it seems by putting snapshots of the spectrogram at each time into such a model, you are actually providing a much better feature.</p>\n<p>On the learning the features side, I've wondered about using something like <a href=\"https://vsitzmann.github.io/siren/\" target=\"_blank\">SIREN</a> to find a data representation - or rather the initial layers of a NN and the initialization method (the latter being one of the big tricks they came up with in the SIREN paper). It seems to me like this and similar ideas for getting cyclic activation functions to work could be a good idea for audio data in principle. The idea is that neural networks could potentially learn something similar to a spectrogram (just better and more problem specific). Of course, all that flexibility might require regularization and tons of data. I.e. I am kind of doubtful about it for the purposes of this competition.</p>\n<p>Also note: There was a <a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\" target=\"_blank\">whole set of TDS articles on this</a> recently (okay, I just realized that was almost 3 years ago…) arguing your viewpoint.</p>",
      "rawMarkdown": "I guess, we kind of know that (Mel-)spectrogram are pretty good features. If you know that, it can in a limited data regimen be preferable to use those vs. learning features. You might wonder about just putting raw signal into a LSTM or some kind of transformer, but it seems by putting snapshots of the spectrogram at each time into such a model, you are actually providing a much better feature.\n\nOn the learning the features side, I've wondered about using something like [SIREN](https://vsitzmann.github.io/siren/) to find a data representation - or rather the initial layers of a NN and the initialization method (the latter being one of the big tricks they came up with in the SIREN paper). It seems to me like this and similar ideas for getting cyclic activation functions to work could be a good idea for audio data in principle. The idea is that neural networks could potentially learn something similar to a spectrogram (just better and more problem specific). Of course, all that flexibility might require regularization and tons of data. I.e. I am kind of doubtful about it for the purposes of this competition.\n\nAlso note: There was a [whole set of TDS articles on this](https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd) recently (okay, I just realized that was almost 3 years ago...) arguing your viewpoint.",
      "votes": null
    },
    {
      "id": "1120617",
      "postDate": "12/21/2020 01:36:53",
      "content": "<p>Thank you for the great references!<br>\nDo you know, if someone already tried to use NNs to produce an output like a spectrogram?</p>",
      "rawMarkdown": "Thank you for the great references!\nDo you know, if someone already tried to use NNs to produce an output like a spectrogram?",
      "votes": null
    },
    {
      "id": "1120897",
      "postDate": "12/21/2020 07:19:39",
      "content": "<p>That sounds like something that should work, so someone might well have tried it, but I'm not aware of such a paper (does of course not mean much given the firehose of deep learning papers we are getting).</p>",
      "rawMarkdown": "That sounds like something that should work, so someone might well have tried it, but I'm not aware of such a paper (does of course not mean much given the firehose of deep learning papers we are getting).",
      "votes": null
    },
    {
      "id": "1121022",
      "postDate": "12/21/2020 09:40:59",
      "content": "<p>I haven't tried it myself, but WaveNet is a good candidate if you want to use raw audio directly</p>",
      "rawMarkdown": "I haven't tried it myself, but WaveNet is a good candidate if you want to use raw audio directly",
      "votes": null
    },
    {
      "id": "1121626",
      "postDate": "12/21/2020 19:22:20",
      "content": "<p>Thanks for the reference. It is also interesting, but isn't that more for generating sound than to anayze sound?</p>",
      "rawMarkdown": "Thanks for the reference. It is also interesting, but isn't that more for generating sound than to anayze sound?",
      "votes": null
    },
    {
      "id": "1121638",
      "postDate": "12/21/2020 19:29:21",
      "content": "<p>I thought about this again and maybe using spectrograms to feed to NNs is not really so different from what happens in our ears, right?<br>\nI mean the cochlea also kind of splits frequencies to make sure different frequencies activate different groups of neurons.</p>",
      "rawMarkdown": "I thought about this again and maybe using spectrograms to feed to NNs is not really so different from what happens in our ears, right?\nI mean the cochlea also kind of splits frequencies to make sure different frequencies activate different groups of neurons.",
      "votes": null
    },
    {
      "id": "1121774",
      "postDate": "12/21/2020 22:34:34",
      "content": "<p>Yeah,the dilated convolutions and lstm are really what make it shine, which is why the SeriesNet adaptation is supposed to work well</p>",
      "rawMarkdown": "Yeah,the dilated convolutions and lstm are really what make it shine, which is why the SeriesNet adaptation is supposed to work well",
      "votes": null
    },
    {
      "id": "1121777",
      "postDate": "12/21/2020 22:36:40",
      "content": "<p>torchlibrosa</p>",
      "rawMarkdown": "torchlibrosa",
      "votes": null
    },
    {
      "id": "1121808",
      "postDate": "12/21/2020 23:27:25",
      "content": "<p>Isn't that just accelerating spectrograms (and other things librosa) does using PyTorch/GPU?</p>",
      "rawMarkdown": "Isn't that just accelerating spectrograms (and other things librosa) does using PyTorch/GPU?",
      "votes": null
    },
    {
      "id": "1121859",
      "postDate": "12/22/2020 01:16:53",
      "content": "<p><a href=\"https://www.kaggle.com/jottbe\" target=\"_blank\">@jottbe</a> You can use wavenet as classifier too. Example here in previous competition : <a href=\"https://www.kaggle.com/siavrez/wavenet-keras\" target=\"_blank\">https://www.kaggle.com/siavrez/wavenet-keras</a></p>",
      "rawMarkdown": "jottbe You can use wavenet as classifier too. Example here in previous competition : https://www.kaggle.com/siavrez/wavenet-keras",
      "votes": null
    },
    {
      "id": "1122100",
      "postDate": "12/22/2020 07:32:21",
      "content": "<p>Final solution 💯 working :-</p>\n<p>This problem was faced by me too finally i have found a solution to this problem with the same approach and model</p>\n<p>Go to this notebook and download the output signals in .npy format </p>\n<p><a href=\"https://www.kaggle.com/tunguz/giba-s-fft-features-only\" target=\"_blank\">https://www.kaggle.com/tunguz/giba-s-fft-features-only</a></p>\n<p>they contain all the signals </p>\n<p>read them in this way </p>\n<p><code>np.load(filename)</code></p>\n<p>train tab are the target variables </p>\n<p>Convert them to spectograms using matplotlib or scipy however this will take a lot time do this on your local computer </p>",
      "rawMarkdown": "Final solution 💯 working :-\n\nThis problem was faced by me too finally i have found a solution to this problem with the same approach and model\n\nGo to this notebook and download the output signals in .npy format \n\nhttps://www.kaggle.com/tunguz/giba-s-fft-features-only\n\nthey contain all the signals \n\nread them in this way \n\n`np.load(filename)`\n\ntrain tab are the target variables \n\nConvert them to spectograms using matplotlib or scipy however this will take a lot time do this on your local computer",
      "votes": null
    },
    {
      "id": "1122939",
      "postDate": "12/22/2020 19:53:42",
      "content": "<p>Thank you for the reference. I'll have a look.</p>",
      "rawMarkdown": "Thank you for the reference. I'll have a look.",
      "votes": null
    },
    {
      "id": "1122946",
      "postDate": "12/22/2020 20:07:34",
      "content": "<p>Thank you for shring this. Do you already have a working model? what is the score of this approach?<br>\nI have seen the score of giba, but I am not sure if the score based on a logistic regression model can tell much about the quality of the features, because it probably doesn't use the full potential.</p>",
      "rawMarkdown": "Thank you for shring this. Do you already have a working model? what is the score of this approach?\nI have seen the score of giba, but I am not sure if the score based on a logistic regression model can tell much about the quality of the features, because it probably doesn't use the full potential.",
      "votes": null
    },
    {
      "id": "1123344",
      "postDate": "12/23/2020 06:42:46",
      "content": "<p>I dont have the model spectogram generation takes time giba used various models not only logit regression </p>",
      "rawMarkdown": "I dont have the model spectogram generation takes time giba used various models not only logit regression",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1120599,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/21/2020 00:55:45",
      "content": "<p>I guess, we kind of know that (Mel-)spectrogram are pretty good features. If you know that, it can in a limited data regimen be preferable to use those vs. learning features. You might wonder about just putting raw signal into a LSTM or some kind of transformer, but it seems by putting snapshots of the spectrogram at each time into such a model, you are actually providing a much better feature.</p>\n<p>On the learning the features side, I've wondered about using something like <a href=\"https://vsitzmann.github.io/siren/\" target=\"_blank\">SIREN</a> to find a data representation - or rather the initial layers of a NN and the initialization method (the latter being one of the big tricks they came up with in the SIREN paper). It seems to me like this and similar ideas for getting cyclic activation functions to work could be a good idea for audio data in principle. The idea is that neural networks could potentially learn something similar to a spectrogram (just better and more problem specific). Of course, all that flexibility might require regularization and tons of data. I.e. I am kind of doubtful about it for the purposes of this competition.</p>\n<p>Also note: There was a <a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\" target=\"_blank\">whole set of TDS articles on this</a> recently (okay, I just realized that was almost 3 years ago…) arguing your viewpoint.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1120617,
          "author_name": "jottbe",
          "author_url": "",
          "post_date": "12/21/2020 01:36:53",
          "content": "<p>Thank you for the great references!<br>\nDo you know, if someone already tried to use NNs to produce an output like a spectrogram?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120897,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/21/2020 07:19:39",
          "content": "<p>That sounds like something that should work, so someone might well have tried it, but I'm not aware of such a paper (does of course not mean much given the firehose of deep learning papers we are getting).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121777,
          "author_name": "eladwar",
          "author_url": "",
          "post_date": "12/21/2020 22:36:40",
          "content": "<p>torchlibrosa</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121808,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/21/2020 23:27:25",
          "content": "<p>Isn't that just accelerating spectrograms (and other things librosa) does using PyTorch/GPU?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121022,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "12/21/2020 09:40:59",
      "content": "<p>I haven't tried it myself, but WaveNet is a good candidate if you want to use raw audio directly</p>",
      "votes": null,
      "replies": [
        {
          "id": 1121626,
          "author_name": "jottbe",
          "author_url": "",
          "post_date": "12/21/2020 19:22:20",
          "content": "<p>Thanks for the reference. It is also interesting, but isn't that more for generating sound than to anayze sound?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121774,
          "author_name": "eladwar",
          "author_url": "",
          "post_date": "12/21/2020 22:34:34",
          "content": "<p>Yeah,the dilated convolutions and lstm are really what make it shine, which is why the SeriesNet adaptation is supposed to work well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121859,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "12/22/2020 01:16:53",
          "content": "<p><a href=\"https://www.kaggle.com/jottbe\" target=\"_blank\">@jottbe</a> You can use wavenet as classifier too. Example here in previous competition : <a href=\"https://www.kaggle.com/siavrez/wavenet-keras\" target=\"_blank\">https://www.kaggle.com/siavrez/wavenet-keras</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122939,
          "author_name": "jottbe",
          "author_url": "",
          "post_date": "12/22/2020 19:53:42",
          "content": "<p>Thank you for the reference. I'll have a look.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121638,
      "author_name": "jottbe",
      "author_url": "",
      "post_date": "12/21/2020 19:29:21",
      "content": "<p>I thought about this again and maybe using spectrograms to feed to NNs is not really so different from what happens in our ears, right?<br>\nI mean the cochlea also kind of splits frequencies to make sure different frequencies activate different groups of neurons.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122100,
      "author_name": "swaralipibose",
      "author_url": "",
      "post_date": "12/22/2020 07:32:21",
      "content": "<p>Final solution 💯 working :-</p>\n<p>This problem was faced by me too finally i have found a solution to this problem with the same approach and model</p>\n<p>Go to this notebook and download the output signals in .npy format </p>\n<p><a href=\"https://www.kaggle.com/tunguz/giba-s-fft-features-only\" target=\"_blank\">https://www.kaggle.com/tunguz/giba-s-fft-features-only</a></p>\n<p>they contain all the signals </p>\n<p>read them in this way </p>\n<p><code>np.load(filename)</code></p>\n<p>train tab are the target variables </p>\n<p>Convert them to spectograms using matplotlib or scipy however this will take a lot time do this on your local computer </p>",
      "votes": null,
      "replies": [
        {
          "id": 1122946,
          "author_name": "jottbe",
          "author_url": "",
          "post_date": "12/22/2020 20:07:34",
          "content": "<p>Thank you for shring this. Do you already have a working model? what is the score of this approach?<br>\nI have seen the score of giba, but I am not sure if the score based on a logistic regression model can tell much about the quality of the features, because it probably doesn't use the full potential.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123344,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "12/23/2020 06:42:46",
          "content": "<p>I dont have the model spectogram generation takes time giba used various models not only logit regression </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1120253": "Hi,\n\nfirst I have to say, that this is my first accoustic data project, so maybe my questions sound stupid to you.\nTo me it seems a bit odd, that we have to produce spectrograms to feed them into a CNN. Ok I know, this is maybe the way a human actually would try to identify a species while hearing to the soud it produces, but don't we really waste a lot of info this way?\n\nWouldn't it be practical to directly feed the audio data (probably after converting it to wav) to a CNN?\nI am sure this was already tried a lot of times, because it seems straightforward to me, but the more I wonder why this method has been abandomed by most of the people working on accoustic signals.\n\nCan anybody shed some light please?\n\nThank you in advance\nJürgen",
    "1120599": "I guess, we kind of know that (Mel-)spectrogram are pretty good features. If you know that, it can in a limited data regimen be preferable to use those vs. learning features. You might wonder about just putting raw signal into a LSTM or some kind of transformer, but it seems by putting snapshots of the spectrogram at each time into such a model, you are actually providing a much better feature.\n\nOn the learning the features side, I've wondered about using something like [SIREN](https://vsitzmann.github.io/siren/) to find a data representation - or rather the initial layers of a NN and the initialization method (the latter being one of the big tricks they came up with in the SIREN paper). It seems to me like this and similar ideas for getting cyclic activation functions to work could be a good idea for audio data in principle. The idea is that neural networks could potentially learn something similar to a spectrogram (just better and more problem specific). Of course, all that flexibility might require regularization and tons of data. I.e. I am kind of doubtful about it for the purposes of this competition.\n\nAlso note: There was a [whole set of TDS articles on this](https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd) recently (okay, I just realized that was almost 3 years ago...) arguing your viewpoint.",
    "1120617": "Thank you for the great references!\nDo you know, if someone already tried to use NNs to produce an output like a spectrogram?",
    "1120897": "That sounds like something that should work, so someone might well have tried it, but I'm not aware of such a paper (does of course not mean much given the firehose of deep learning papers we are getting).",
    "1121022": "I haven't tried it myself, but WaveNet is a good candidate if you want to use raw audio directly",
    "1121626": "Thanks for the reference. It is also interesting, but isn't that more for generating sound than to anayze sound?",
    "1121638": "I thought about this again and maybe using spectrograms to feed to NNs is not really so different from what happens in our ears, right?\nI mean the cochlea also kind of splits frequencies to make sure different frequencies activate different groups of neurons.",
    "1121774": "Yeah,the dilated convolutions and lstm are really what make it shine, which is why the SeriesNet adaptation is supposed to work well",
    "1121777": "torchlibrosa",
    "1121808": "Isn't that just accelerating spectrograms (and other things librosa) does using PyTorch/GPU?",
    "1121859": "jottbe You can use wavenet as classifier too. Example here in previous competition : https://www.kaggle.com/siavrez/wavenet-keras",
    "1122100": "Final solution 💯 working :-\n\nThis problem was faced by me too finally i have found a solution to this problem with the same approach and model\n\nGo to this notebook and download the output signals in .npy format \n\nhttps://www.kaggle.com/tunguz/giba-s-fft-features-only\n\nthey contain all the signals \n\nread them in this way \n\n`np.load(filename)`\n\ntrain tab are the target variables \n\nConvert them to spectograms using matplotlib or scipy however this will take a lot time do this on your local computer",
    "1122939": "Thank you for the reference. I'll have a look.",
    "1122946": "Thank you for shring this. Do you already have a working model? what is the score of this approach?\nI have seen the score of giba, but I am not sure if the score based on a logistic regression model can tell much about the quality of the features, because it probably doesn't use the full potential.",
    "1123344": "I dont have the model spectogram generation takes time giba used various models not only logit regression"
  },
  "source": "meta"
}