{
  "id": 161505,
  "title": "Anyone Try WaveNet?",
  "url": "/competitions/birdsong-recognition/discussion/161505",
  "author_name": "",
  "post_date": "2020-06-25T05:09:38.636072200Z",
  "votes": 29,
  "comment_count": 17,
  "views": 0,
  "content": "<p>There are many examples of TensorFlow and PyTorch code for WaveNet in Ion Comp notebooks. Has anyone tried applying WaveNet for classifying birdcalls?</p>\n\n<p>I suggest using normal convolutions (not causal). Do 4 wavenet blocks and then add a classification head on top. TensorFlow WaveNet <a href=\"https://www.kaggle.com/siavrez/wavenet-keras\">here</a>. PyTorch WaveNet <a href=\"https://www.kaggle.com/cswwp347724/wavenet-pytorch\">here</a></p>\n\n<p>This would allow you to leave the train data as 1D and not convert it to 2D spectrograms. WaveNet does a good job at identifying all the component frequencies similar to a FFT spectrogram.</p>",
  "messages": [
    {
      "id": "900839",
      "postDate": "06/25/2020 05:09:38",
      "content": "<p>There are many examples of TensorFlow and PyTorch code for WaveNet in Ion Comp notebooks. Has anyone tried applying WaveNet for classifying birdcalls?</p>\n\n<p>I suggest using normal convolutions (not causal). Do 4 wavenet blocks and then add a classification head on top. TensorFlow WaveNet <a href=\"https://www.kaggle.com/siavrez/wavenet-keras\">here</a>. PyTorch WaveNet <a href=\"https://www.kaggle.com/cswwp347724/wavenet-pytorch\">here</a></p>\n\n<p>This would allow you to leave the train data as 1D and not convert it to 2D spectrograms. WaveNet does a good job at identifying all the component frequencies similar to a FFT spectrogram.</p>",
      "rawMarkdown": "There are many examples of TensorFlow and PyTorch code for WaveNet in Ion Comp notebooks. Has anyone tried applying WaveNet for classifying birdcalls?\n\nI suggest using normal convolutions (not causal). Do 4 wavenet blocks and then add a classification head on top. TensorFlow WaveNet [here][1]. PyTorch WaveNet [here][2]\n\nThis would allow you to leave the train data as 1D and not convert it to 2D spectrograms. WaveNet does a good job at identifying all the component frequencies similar to a FFT spectrogram.\n\n[1]: https://www.kaggle.com/siavrez/wavenet-keras\n[2]: https://www.kaggle.com/cswwp347724/wavenet-pytorch",
      "votes": null
    },
    {
      "id": "902628",
      "postDate": "06/26/2020 09:15:43",
      "content": "<p>Hi, Thank you share idea.\nI'm trying it now.</p>\n\n<ul>\n<li>Dataset: <a href=\"https://www.kaggle.com/takamichitoda/birdcall-dataset-for-wavenet\">https://www.kaggle.com/takamichitoda/birdcall-dataset-for-wavenet</a></li>\n<li>Notebook: <a href=\"https://www.kaggle.com/takamichitoda/birdcall-wavenet-very-very-small-baseline?scriptVersionId=37468713\">https://www.kaggle.com/takamichitoda/birdcall-wavenet-very-very-small-baseline?scriptVersionId=37468713</a></li>\n</ul>\n\n<p>However, I reach GPU usage limit.(kaggle notebook and google colab)\nStarting tomorrow, I will try to organize the code and try various ideas.</p>",
      "rawMarkdown": "Hi, Thank you share idea.\nI'm trying it now.\n\n- Dataset: https://www.kaggle.com/takamichitoda/birdcall-dataset-for-wavenet\n- Notebook: https://www.kaggle.com/takamichitoda/birdcall-wavenet-very-very-small-baseline?scriptVersionId=37468713\n\nHowever, I reach GPU usage limit.(kaggle notebook and google colab)\nStarting tomorrow, I will try to organize the code and try various ideas.",
      "votes": null
    },
    {
      "id": "902810",
      "postDate": "06/26/2020 11:27:23",
      "content": "<p>Great suggestion! I would definitely try implementing it. </p>\n\n<p>Could you please review <a href=\"https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe?rvi=1\">my notebook</a> on feature extraction of the audio data? It was my first detailed study of audio data and an advice would be really very helpful. Thank You :)</p>",
      "rawMarkdown": "Great suggestion! I would definitely try implementing it. \n\nCould you please review [my notebook](https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe?rvi=1) on feature extraction of the audio data? It was my first detailed study of audio data and an advice would be really very helpful. Thank You :)",
      "votes": null
    },
    {
      "id": "907749",
      "postDate": "06/30/2020 06:28:32",
      "content": "<p>I have done some experiment, and I share the gotten little insight.</p>\n\n<p>It is not recommended to use pure raw audio data by WaveNet, because it has many noise.\nI think that μ-law algorithm used by WaveNet increase noise.</p>\n\n<p>If we use WaveNet, we should do denoise.\nI will try denoise(but I don't have good idea now)</p>\n\n<p>I'm not specialist of audio, if my insight is mistake, please comment.</p>",
      "rawMarkdown": "I have done some experiment, and I share the gotten little insight.\n\nIt is not recommended to use pure raw audio data by WaveNet, because it has many noise.\nI think that μ-law algorithm used by WaveNet increase noise.\n\nIf we use WaveNet, we should do denoise.\nI will try denoise(but I don't have good idea now)\n\nI'm not specialist of audio, if my insight is mistake, please comment.",
      "votes": null
    },
    {
      "id": "909911",
      "postDate": "06/30/2020 23:17:57",
      "content": "<p>I just plan to use WaveNet since it performed very well and lead us to silver medal (which we have not chosen as submission). It would be nice to see some implementation here. </p>",
      "rawMarkdown": "I just plan to use WaveNet since it performed very well and lead us to silver medal (which we have not chosen as submission). It would be nice to see some implementation here.",
      "votes": null
    },
    {
      "id": "911732",
      "postDate": "07/02/2020 01:48:02",
      "content": "<p>I agree. WaveNet did awesome in Ion Comp, so I'm eager to try it here.</p>",
      "rawMarkdown": "I agree. WaveNet did awesome in Ion Comp, so I'm eager to try it here.",
      "votes": null
    },
    {
      "id": "911733",
      "postDate": "07/02/2020 01:49:03",
      "content": "<p>Thanks for the suggestions. I plan to try WaveNet soon. I will be careful to remove noise.</p>",
      "rawMarkdown": "Thanks for the suggestions. I plan to try WaveNet soon. I will be careful to remove noise.",
      "votes": null
    },
    {
      "id": "911734",
      "postDate": "07/02/2020 01:49:59",
      "content": "<p>Great notebook!</p>",
      "rawMarkdown": "Great notebook!",
      "votes": null
    },
    {
      "id": "911737",
      "postDate": "07/02/2020 01:52:10",
      "content": "<p>Going to try it tomorrow, was making the preprocess script to transform the audio data into tensors </p>",
      "rawMarkdown": "Going to try it tomorrow, was making the preprocess script to transform the audio data into tensors",
      "votes": null
    },
    {
      "id": "929101",
      "postDate": "07/14/2020 13:10:17",
      "content": "<p>Thx for sharing.\nI read <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253\">your post in Ion Comp</a>, I found that you used raw audio signal as the input of WaveNet. \nIs the raw signal better than mfcc or fb40 when using WaveNet? \nI used to use fb40 with some convolution networks like vgg or something else, I'm not familiar with WaveNet.</p>",
      "rawMarkdown": "Thx for sharing.\nI read [your post in Ion Comp](https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253), I found that you used raw audio signal as the input of WaveNet. \nIs the raw signal better than mfcc or fb40 when using WaveNet? \nI used to use fb40 with some convolution networks like vgg or something else, I'm not familiar with WaveNet.",
      "votes": null
    },
    {
      "id": "929402",
      "postDate": "07/14/2020 16:43:49",
      "content": "<p>The 'generalized' form of WaveNet is the Temporal Convolutional Network:\n<a href=\"https://arxiv.org/abs/1608.08242\">https://arxiv.org/abs/1608.08242</a></p>\n\n<p>You can use it on raw audio or on spectrograms (or mel spectrograms or mfccs or whatever you like); the core idea is stacking dilated convolutions to aggregate over large time windows, which can work in any domain. Presumably, a well-crafted melspectrogram should have much lower dimensionality than raw audio, so using melspecs should be a bit more memory efficient for looking at very long time spans, but also throws out some information in the process of building the melspec....</p>\n\n<p>So, I'm not sure which will perform better, but I'll be very happy to hear the results if someone tries both!</p>",
      "rawMarkdown": "The 'generalized' form of WaveNet is the Temporal Convolutional Network:\nhttps://arxiv.org/abs/1608.08242\n\nYou can use it on raw audio or on spectrograms (or mel spectrograms or mfccs or whatever you like); the core idea is stacking dilated convolutions to aggregate over large time windows, which can work in any domain. Presumably, a well-crafted melspectrogram should have much lower dimensionality than raw audio, so using melspecs should be a bit more memory efficient for looking at very long time spans, but also throws out some information in the process of building the melspec....\n\nSo, I'm not sure which will perform better, but I'll be very happy to hear the results if someone tries both!",
      "votes": null
    },
    {
      "id": "929842",
      "postDate": "07/15/2020 02:42:52",
      "content": "<p>Thanks for replying. I will try all of them.</p>\n\n<p>As <a href=\"/cdeotte\">@cdeotte</a> says <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253\">in this post</a> </p>\n\n<blockquote>\n  <p>WaveNet is superior at finding patterns in waveforms. Using dilated convolutions, it decomposes a signal into its different frequency sine waves just like the Fourier Transform.</p>\n</blockquote>\n\n<p>So I have this question is raw signal the best input for wavnet.</p>",
      "rawMarkdown": "Thanks for replying. I will try all of them.\n\nAs @cdeotte says [in this post](https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253) \n\n&gt; WaveNet is superior at finding patterns in waveforms. Using dilated convolutions, it decomposes a signal into its different frequency sine waves just like the Fourier Transform.\n\nSo I have this question is raw signal the best input for wavnet.",
      "votes": null
    },
    {
      "id": "950452",
      "postDate": "07/29/2020 12:03:33",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Any plans to release a tensorflow kernel here?🤒 </p>",
      "rawMarkdown": "cdeotte Any plans to release a tensorflow kernel here?🤒",
      "votes": null
    },
    {
      "id": "954637",
      "postDate": "08/02/2020 00:12:43",
      "content": "<p>Probably not soon. I do love audio and I'm interested in joining this comp but I've been busy with other things.</p>",
      "rawMarkdown": "Probably not soon. I do love audio and I'm interested in joining this comp but I've been busy with other things.",
      "votes": null
    },
    {
      "id": "961237",
      "postDate": "08/07/2020 03:08:40",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  The calls are short, the songs are longer. How can we decide the batch length in this case (in Ion, 4000 was used and it was a sequence to sequence prediction, unlike here)<br>\nThanks in advance if you have some hints. I hope you will be less busy soon :) </p>",
      "rawMarkdown": "cdeotte  The calls are short, the songs are longer. How can we decide the batch length in this case (in Ion, 4000 was used and it was a sequence to sequence prediction, unlike here)\nThanks in advance if you have some hints. I hope you will be less busy soon :)",
      "votes": null
    },
    {
      "id": "961243",
      "postDate": "08/07/2020 03:17:36",
      "content": "<p>Based on the sampling rate in Ion comp which was 10kHz, using 4000 time steps in Ion comp with dilated convolutions of <code>[2**i for i in range(12)]</code> allowed us to analyze all frequencies between about 5Hz to 10kHz. I'm not an expert on WaveNet, but I would try picking a time window and dilation window that covers the relevant frequencies.</p>",
      "rawMarkdown": "Based on the sampling rate in Ion comp which was 10kHz, using 4000 time steps in Ion comp with dilated convolutions of `[2**i for i in range(12)]` allowed us to analyze all frequencies between about 5Hz to 10kHz. I'm not an expert on WaveNet, but I would try picking a time window and dilation window that covers the relevant frequencies.",
      "votes": null
    },
    {
      "id": "961244",
      "postDate": "08/07/2020 03:19:10",
      "content": "<p>And then you would need to \"slide\" this time window over the entire sound clip and use all that information to classify.</p>",
      "rawMarkdown": "And then you would need to \"slide\" this time window over the entire sound clip and use all that information to classify.",
      "votes": null
    },
    {
      "id": "961548",
      "postDate": "08/07/2020 09:23:51",
      "content": "<p>I try denoise this notebook.\n<a href=\"https://www.kaggle.com/takamichitoda/birdcall-noise-reduction\">https://www.kaggle.com/takamichitoda/birdcall-noise-reduction</a></p>\n\n<p>The denoise looks succeed, but not good work in training spectrogram.\nI will re-try training WaveNet.</p>",
      "rawMarkdown": "I try denoise this notebook.\nhttps://www.kaggle.com/takamichitoda/birdcall-noise-reduction\n\nThe denoise looks succeed, but not good work in training spectrogram.\nI will re-try training WaveNet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 902628,
      "author_name": "takamichitoda",
      "author_url": "",
      "post_date": "06/26/2020 09:15:43",
      "content": "<p>Hi, Thank you share idea.\nI'm trying it now.</p>\n\n<ul>\n<li>Dataset: <a href=\"https://www.kaggle.com/takamichitoda/birdcall-dataset-for-wavenet\">https://www.kaggle.com/takamichitoda/birdcall-dataset-for-wavenet</a></li>\n<li>Notebook: <a href=\"https://www.kaggle.com/takamichitoda/birdcall-wavenet-very-very-small-baseline?scriptVersionId=37468713\">https://www.kaggle.com/takamichitoda/birdcall-wavenet-very-very-small-baseline?scriptVersionId=37468713</a></li>\n</ul>\n\n<p>However, I reach GPU usage limit.(kaggle notebook and google colab)\nStarting tomorrow, I will try to organize the code and try various ideas.</p>",
      "votes": null,
      "replies": [
        {
          "id": 907749,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "06/30/2020 06:28:32",
          "content": "<p>I have done some experiment, and I share the gotten little insight.</p>\n\n<p>It is not recommended to use pure raw audio data by WaveNet, because it has many noise.\nI think that μ-law algorithm used by WaveNet increase noise.</p>\n\n<p>If we use WaveNet, we should do denoise.\nI will try denoise(but I don't have good idea now)</p>\n\n<p>I'm not specialist of audio, if my insight is mistake, please comment.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911733,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/02/2020 01:49:03",
          "content": "<p>Thanks for the suggestions. I plan to try WaveNet soon. I will be careful to remove noise.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961548,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "08/07/2020 09:23:51",
          "content": "<p>I try denoise this notebook.\n<a href=\"https://www.kaggle.com/takamichitoda/birdcall-noise-reduction\">https://www.kaggle.com/takamichitoda/birdcall-noise-reduction</a></p>\n\n<p>The denoise looks succeed, but not good work in training spectrogram.\nI will re-try training WaveNet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 902810,
      "author_name": "navinmundhra",
      "author_url": "",
      "post_date": "06/26/2020 11:27:23",
      "content": "<p>Great suggestion! I would definitely try implementing it. </p>\n\n<p>Could you please review <a href=\"https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe?rvi=1\">my notebook</a> on feature extraction of the audio data? It was my first detailed study of audio data and an advice would be really very helpful. Thank You :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 911734,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/02/2020 01:49:59",
          "content": "<p>Great notebook!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 909911,
      "author_name": "muhakabartay",
      "author_url": "",
      "post_date": "06/30/2020 23:17:57",
      "content": "<p>I just plan to use WaveNet since it performed very well and lead us to silver medal (which we have not chosen as submission). It would be nice to see some implementation here. </p>",
      "votes": null,
      "replies": [
        {
          "id": 911732,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/02/2020 01:48:02",
          "content": "<p>I agree. WaveNet did awesome in Ion Comp, so I'm eager to try it here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911737,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "07/02/2020 01:52:10",
          "content": "<p>Going to try it tomorrow, was making the preprocess script to transform the audio data into tensors </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 929101,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "07/14/2020 13:10:17",
      "content": "<p>Thx for sharing.\nI read <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253\">your post in Ion Comp</a>, I found that you used raw audio signal as the input of WaveNet. \nIs the raw signal better than mfcc or fb40 when using WaveNet? \nI used to use fb40 with some convolution networks like vgg or something else, I'm not familiar with WaveNet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 929402,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "07/14/2020 16:43:49",
          "content": "<p>The 'generalized' form of WaveNet is the Temporal Convolutional Network:\n<a href=\"https://arxiv.org/abs/1608.08242\">https://arxiv.org/abs/1608.08242</a></p>\n\n<p>You can use it on raw audio or on spectrograms (or mel spectrograms or mfccs or whatever you like); the core idea is stacking dilated convolutions to aggregate over large time windows, which can work in any domain. Presumably, a well-crafted melspectrogram should have much lower dimensionality than raw audio, so using melspecs should be a bit more memory efficient for looking at very long time spans, but also throws out some information in the process of building the melspec....</p>\n\n<p>So, I'm not sure which will perform better, but I'll be very happy to hear the results if someone tries both!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929842,
          "author_name": "karlyukang",
          "author_url": "",
          "post_date": "07/15/2020 02:42:52",
          "content": "<p>Thanks for replying. I will try all of them.</p>\n\n<p>As <a href=\"/cdeotte\">@cdeotte</a> says <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253\">in this post</a> </p>\n\n<blockquote>\n  <p>WaveNet is superior at finding patterns in waveforms. Using dilated convolutions, it decomposes a signal into its different frequency sine waves just like the Fourier Transform.</p>\n</blockquote>\n\n<p>So I have this question is raw signal the best input for wavnet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 950452,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "07/29/2020 12:03:33",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Any plans to release a tensorflow kernel here?🤒 </p>",
      "votes": null,
      "replies": [
        {
          "id": 954637,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/02/2020 00:12:43",
          "content": "<p>Probably not soon. I do love audio and I'm interested in joining this comp but I've been busy with other things.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961237,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "08/07/2020 03:08:40",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  The calls are short, the songs are longer. How can we decide the batch length in this case (in Ion, 4000 was used and it was a sequence to sequence prediction, unlike here)<br>\nThanks in advance if you have some hints. I hope you will be less busy soon :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961243,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/07/2020 03:17:36",
          "content": "<p>Based on the sampling rate in Ion comp which was 10kHz, using 4000 time steps in Ion comp with dilated convolutions of <code>[2**i for i in range(12)]</code> allowed us to analyze all frequencies between about 5Hz to 10kHz. I'm not an expert on WaveNet, but I would try picking a time window and dilation window that covers the relevant frequencies.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961244,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/07/2020 03:19:10",
          "content": "<p>And then you would need to \"slide\" this time window over the entire sound clip and use all that information to classify.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "900839": "There are many examples of TensorFlow and PyTorch code for WaveNet in Ion Comp notebooks. Has anyone tried applying WaveNet for classifying birdcalls?\n\nI suggest using normal convolutions (not causal). Do 4 wavenet blocks and then add a classification head on top. TensorFlow WaveNet [here][1]. PyTorch WaveNet [here][2]\n\nThis would allow you to leave the train data as 1D and not convert it to 2D spectrograms. WaveNet does a good job at identifying all the component frequencies similar to a FFT spectrogram.\n\n[1]: https://www.kaggle.com/siavrez/wavenet-keras\n[2]: https://www.kaggle.com/cswwp347724/wavenet-pytorch",
    "902628": "Hi, Thank you share idea.\nI'm trying it now.\n\n- Dataset: https://www.kaggle.com/takamichitoda/birdcall-dataset-for-wavenet\n- Notebook: https://www.kaggle.com/takamichitoda/birdcall-wavenet-very-very-small-baseline?scriptVersionId=37468713\n\nHowever, I reach GPU usage limit.(kaggle notebook and google colab)\nStarting tomorrow, I will try to organize the code and try various ideas.",
    "902810": "Great suggestion! I would definitely try implementing it. \n\nCould you please review [my notebook](https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe?rvi=1) on feature extraction of the audio data? It was my first detailed study of audio data and an advice would be really very helpful. Thank You :)",
    "907749": "I have done some experiment, and I share the gotten little insight.\n\nIt is not recommended to use pure raw audio data by WaveNet, because it has many noise.\nI think that μ-law algorithm used by WaveNet increase noise.\n\nIf we use WaveNet, we should do denoise.\nI will try denoise(but I don't have good idea now)\n\nI'm not specialist of audio, if my insight is mistake, please comment.",
    "909911": "I just plan to use WaveNet since it performed very well and lead us to silver medal (which we have not chosen as submission). It would be nice to see some implementation here.",
    "911732": "I agree. WaveNet did awesome in Ion Comp, so I'm eager to try it here.",
    "911733": "Thanks for the suggestions. I plan to try WaveNet soon. I will be careful to remove noise.",
    "911734": "Great notebook!",
    "911737": "Going to try it tomorrow, was making the preprocess script to transform the audio data into tensors",
    "929101": "Thx for sharing.\nI read [your post in Ion Comp](https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253), I found that you used raw audio signal as the input of WaveNet. \nIs the raw signal better than mfcc or fb40 when using WaveNet? \nI used to use fb40 with some convolution networks like vgg or something else, I'm not familiar with WaveNet.",
    "929402": "The 'generalized' form of WaveNet is the Temporal Convolutional Network:\nhttps://arxiv.org/abs/1608.08242\n\nYou can use it on raw audio or on spectrograms (or mel spectrograms or mfccs or whatever you like); the core idea is stacking dilated convolutions to aggregate over large time windows, which can work in any domain. Presumably, a well-crafted melspectrogram should have much lower dimensionality than raw audio, so using melspecs should be a bit more memory efficient for looking at very long time spans, but also throws out some information in the process of building the melspec....\n\nSo, I'm not sure which will perform better, but I'll be very happy to hear the results if someone tries both!",
    "929842": "Thanks for replying. I will try all of them.\n\nAs @cdeotte says [in this post](https://www.kaggle.com/c/liverpool-ion-switching/discussion/154253) \n\n&gt; WaveNet is superior at finding patterns in waveforms. Using dilated convolutions, it decomposes a signal into its different frequency sine waves just like the Fourier Transform.\n\nSo I have this question is raw signal the best input for wavnet.",
    "950452": "cdeotte Any plans to release a tensorflow kernel here?🤒",
    "954637": "Probably not soon. I do love audio and I'm interested in joining this comp but I've been busy with other things.",
    "961237": "cdeotte  The calls are short, the songs are longer. How can we decide the batch length in this case (in Ion, 4000 was used and it was a sequence to sequence prediction, unlike here)\nThanks in advance if you have some hints. I hope you will be less busy soon :)",
    "961243": "Based on the sampling rate in Ion comp which was 10kHz, using 4000 time steps in Ion comp with dilated convolutions of `[2**i for i in range(12)]` allowed us to analyze all frequencies between about 5Hz to 10kHz. I'm not an expert on WaveNet, but I would try picking a time window and dilation window that covers the relevant frequencies.",
    "961244": "And then you would need to \"slide\" this time window over the entire sound clip and use all that information to classify.",
    "961548": "I try denoise this notebook.\nhttps://www.kaggle.com/takamichitoda/birdcall-noise-reduction\n\nThe denoise looks succeed, but not good work in training spectrogram.\nI will re-try training WaveNet."
  },
  "source": "meta"
}