{
  "id": 268768,
  "title": "Few things about CQT",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/268768",
  "author_name": "",
  "post_date": "2021-08-28T19:59:52.326980Z",
  "votes": 28,
  "comment_count": 19,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>I have discovered this transformation thanks to the following <a href=\"https://www.kaggle.com/atamazian/nnaudio-constant-q-transform-demonstration\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">Araik Tamazian</a>. </p>\n<p>It is a <strong>signal processing</strong> transformation closely related to the <a href=\"https://en.wikipedia.org/wiki/Fourier_transform\" target=\"_blank\"><strong>Fourier transform</strong></a> (more specifically to the <a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\" target=\"_blank\">short term</a> version of it). </p>\n<p>It is also related to the complex <a href=\"https://en.wikipedia.org/wiki/Morlet_wavelet\" target=\"_blank\">Morlet wavelet</a>. We will see how in a later section.</p>\n<p>Also, notice that the full name is <strong>constant-Q transform</strong> (or CQT in short). This name comes from the fact<br>\nthat the transformation can be thought of as applying different filters equally spaced on a logarithmic frequency <br>\nscale. The k-th filter is related to the (k-1)-th filter using the following formula (from Wikipeda):</p>\n<p><a href=\"https://ibb.co/Wv4nwdf\"><img src=\"https://i.ibb.co/0yRqSWr/delta-fk-cqt.png\" alt=\"delta-fk-cqt\"></a></p>\n<p>where ''δf''_'k'' is the bandwidth of the ''k''-th filter, ''f''_min is the central frequency of the lowest filter, and ''n'' is the number of filters per octave.</p>\n<p>The <strong>CQT</strong> has been used a lot in music since it is well suited to the limited range of human's auditory perception <br>\n(20 to 20k Hz) and to the fact that the human ear is more sensitive to lower frequencies than higher ones.</p>\n<h1>Theoretical computation</h1>\n<p>Here is how it is defined and computed in few steps:</p>\n<ol>\n<li>We start by computing the <a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\" target=\"_blank\">short-term Fourier transform</a> (STFT in short): this is a Fourier transform with a moving window over the signal.<br>\nAs a window function, we often use the <a href=\"https://en.wikipedia.org/wiki/Window_function#Hann_and_Hamming_windows\" target=\"_blank\">Hann or Hamming windows</a>.<br>\nFinally, notice that we will use the discrete version of the STFT. The equation should be: </li>\n</ol>\n<p><a href=\"https://ibb.co/KwNjb7y\"><img src=\"https://i.ibb.co/HdNDBYg/stft-cqt.png\" alt=\"stft-cqt\"></a></p>\n<ol>\n<li>We introduce the <strong>quality factor</strong> <code>Q</code> using the following formula: </li>\n</ol>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/12ZckR8/q-cqt.png\" alt=\"q-cqt\"></a></p>\n<ol>\n<li>Now, the window length for the k-th bin is no longer fixed (N) by will vary depending on the filter <br>\n(we use the previous definition to get a simplified expression): </li>\n</ol>\n<p><a href=\"https://ibb.co/crwqYgc\"><img src=\"https://i.ibb.co/YLXg0R7/n-cqt.png\" alt=\"n-cqt\"></a></p>\n<ol>\n<li><p>The digital frequency \\frac{2 \\pi k}{N} is also changed and the window function now depends on k as well (via the window length N[k])</p></li>\n<li><p>Finally, replacing the different elements in the original formula, we get: </p></li>\n</ol>\n<p><a href=\"https://ibb.co/T48ZfYL\"><img src=\"https://i.ibb.co/wSWTkBs/cqt-cqt.png\" alt=\"cqt-cqt\"></a></p>\n<p>Notice that for a fast implementation, <a href=\"https://en.wikipedia.org/wiki/Fast_Fourier_transform\" target=\"_blank\">FFT</a> is often used. Sometimes, it can even be replaced <br>\nwith a <a href=\"https://en.wikipedia.org/wiki/Sliding_DFT\" target=\"_blank\">sliding DFT</a>.</p>\n<h1>&nbsp;Implementation</h1>\n<p>If you have access to a GPU and are using it with a deep learning <a href=\"https://pytorch.org/\" target=\"_blank\">PyTorch</a> model, the best library to use is probably <code>nnAudio</code> since it gives a huge execution time boost since it runs on GPU. The code follows the algorithm proposed in the following <a href=\"http://academics.wellesley.edu/Physics/brown/pubs/effalgV92P2698-P2701.pdf\" target=\"_blank\">paper</a> with a slight improvement.</p>\n<p>The main function that we will be using is <code>CQT1992v2</code>: <strong>1992</strong> stands for the year the algorithm's paper has been published and <strong>v2</strong> stands for a slight improvement on the original paper. </p>\n<p>Here is a code snippet to do it (notice that  the code below will run on CPU): </p>\n<pre><code>import torch\nfrom nnAudio.Spectrogram import CQT1992v2\nimport matplotlib.pylab as plt\nimport numpy as np\n\n# sr is the sampling rate, it is 2048 Hz\n# fmax is half the sampling rate\n\ncqt_transform = CQT1992v2(sr=2048, fmin=20, fmax=1024, hop_length=64)\n\ndef run_cqt_transform(x: np.array) -&gt; torch.Tensor:\n    # We stack the passed x since there are 3\n    # time series per file.\n    x = np.hstack(x)\n    # Normalize (is there a better way?)\n    x = x / np.max(x)\n    x = torch.from_numpy(x).float()\n    return cqt_transform(x)\n\n# Running on one file and plotting the result.\nx = np.load(\"path/to/file.npy\")\n\n# We take the first (and only) result since the \n# result is batch-shaped ((1, freq_bins, time_steps)).\nimg = run_cqt_transform(x)[0]\n\nfig, ax = plt.subplots(1, 1, figsize=(12, 8))\nax.imshow(img)\nplt.show()\n</code></pre>\n<p>Here is a sample of an output that you can get:</p>\n<p><a href=\"https://ibb.co/tDHZ5Ph\"><img src=\"https://i.ibb.co/5L1YVx9/cqt-sample-output.png\" alt=\"cqt-sample-output\"></a></p>\n<p>Notice that there are many other parameters that you can pass to <code>CQT1992v2</code>. <br>\nOne of these is the <code>bins_per_octave</code> which defaults to <code>12</code> and the <code>output_format</code> which defaults to <code>Magnitude</code> (i.e. the complex module). You can also get the real and imaginary parts with <code>Complex</code>.</p>\n<h1>Link to the Morlet wavelet transform</h1>\n<p>If we use the <a href=\"https://en.wikipedia.org/wiki/Morlet_wavelet\" target=\"_blank\">Morlet wavelet</a>, i.e. a windowed (complex) sinusoidal where the window is a Gaussian, in a continuous wavelet transform it becomes a kind of STFT with the Morelet wavelet replacing the window part X[n-m] times the (complex) sinusoidal. Then, doing the same steps as in Theoretical computation section, we get<br>\nan analogous constant-q transform (the filters have the correct scaling thanks to the wavelet's scaling). </p>\n<p>The advantage of the Morlet wavelet transform in comparison with CQT is that its<br>\ninverse is easier to compute since the wavelets form a basis, whereas in CQT it isn't the case in general.</p>\n<p>More details can be found <a href=\"https://ccrma.stanford.edu/~jos/sasp/Continuous_Wavelet_Transform.html\" target=\"_blank\">here</a>.</p>\n<h1>Resources</h1>\n<p>For more details, check the Wikipedia <a href=\"https://en.wikipedia.org/wiki/Constant-Q_transform\" target=\"_blank\">page</a>.</p>\n<p>And finally, for those that like to check even more things, here are links to related concepts and further explanations:</p>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Fourier_transform\" target=\"_blank\">Fourier transform</a></li>\n<li><a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\" target=\"_blank\">Short-time Fourier transform</a></li>\n<li><a href=\"https://en.wikipedia.org/wiki/Gabor_transform\" target=\"_blank\">Gabor transform</a></li>\n<li><a href=\"https://en.wikipedia.org/wiki/List_of_Fourier-related_transforms\" target=\"_blank\">Many more Fourier related transforms</a></li>\n<li>A <a href=\"https://www.youtube.com/watch?v=EfWnEldTyPA\" target=\"_blank\">video</a> about the spectrogram and the Gabor transform </li>\n<li>A <a href=\"https://www.youtube.com/watch?v=Cl-m4X3rwac\" target=\"_blank\">video</a> about the constant-q transform</li>\n</ul>\n<p>[UPDATE] Few small code fixes thanks to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> comment. </p>",
  "messages": [
    {
      "id": "1494588",
      "postDate": "08/28/2021 19:59:52",
      "content": "<h1>Introduction</h1>\n<p>I have discovered this transformation thanks to the following <a href=\"https://www.kaggle.com/atamazian/nnaudio-constant-q-transform-demonstration\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">Araik Tamazian</a>. </p>\n<p>It is a <strong>signal processing</strong> transformation closely related to the <a href=\"https://en.wikipedia.org/wiki/Fourier_transform\" target=\"_blank\"><strong>Fourier transform</strong></a> (more specifically to the <a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\" target=\"_blank\">short term</a> version of it). </p>\n<p>It is also related to the complex <a href=\"https://en.wikipedia.org/wiki/Morlet_wavelet\" target=\"_blank\">Morlet wavelet</a>. We will see how in a later section.</p>\n<p>Also, notice that the full name is <strong>constant-Q transform</strong> (or CQT in short). This name comes from the fact<br>\nthat the transformation can be thought of as applying different filters equally spaced on a logarithmic frequency <br>\nscale. The k-th filter is related to the (k-1)-th filter using the following formula (from Wikipeda):</p>\n<p><a href=\"https://ibb.co/Wv4nwdf\"><img src=\"https://i.ibb.co/0yRqSWr/delta-fk-cqt.png\" alt=\"delta-fk-cqt\"></a></p>\n<p>where ''δf''_'k'' is the bandwidth of the ''k''-th filter, ''f''_min is the central frequency of the lowest filter, and ''n'' is the number of filters per octave.</p>\n<p>The <strong>CQT</strong> has been used a lot in music since it is well suited to the limited range of human's auditory perception <br>\n(20 to 20k Hz) and to the fact that the human ear is more sensitive to lower frequencies than higher ones.</p>\n<h1>Theoretical computation</h1>\n<p>Here is how it is defined and computed in few steps:</p>\n<ol>\n<li>We start by computing the <a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\" target=\"_blank\">short-term Fourier transform</a> (STFT in short): this is a Fourier transform with a moving window over the signal.<br>\nAs a window function, we often use the <a href=\"https://en.wikipedia.org/wiki/Window_function#Hann_and_Hamming_windows\" target=\"_blank\">Hann or Hamming windows</a>.<br>\nFinally, notice that we will use the discrete version of the STFT. The equation should be: </li>\n</ol>\n<p><a href=\"https://ibb.co/KwNjb7y\"><img src=\"https://i.ibb.co/HdNDBYg/stft-cqt.png\" alt=\"stft-cqt\"></a></p>\n<ol>\n<li>We introduce the <strong>quality factor</strong> <code>Q</code> using the following formula: </li>\n</ol>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/12ZckR8/q-cqt.png\" alt=\"q-cqt\"></a></p>\n<ol>\n<li>Now, the window length for the k-th bin is no longer fixed (N) by will vary depending on the filter <br>\n(we use the previous definition to get a simplified expression): </li>\n</ol>\n<p><a href=\"https://ibb.co/crwqYgc\"><img src=\"https://i.ibb.co/YLXg0R7/n-cqt.png\" alt=\"n-cqt\"></a></p>\n<ol>\n<li><p>The digital frequency \\frac{2 \\pi k}{N} is also changed and the window function now depends on k as well (via the window length N[k])</p></li>\n<li><p>Finally, replacing the different elements in the original formula, we get: </p></li>\n</ol>\n<p><a href=\"https://ibb.co/T48ZfYL\"><img src=\"https://i.ibb.co/wSWTkBs/cqt-cqt.png\" alt=\"cqt-cqt\"></a></p>\n<p>Notice that for a fast implementation, <a href=\"https://en.wikipedia.org/wiki/Fast_Fourier_transform\" target=\"_blank\">FFT</a> is often used. Sometimes, it can even be replaced <br>\nwith a <a href=\"https://en.wikipedia.org/wiki/Sliding_DFT\" target=\"_blank\">sliding DFT</a>.</p>\n<h1>&nbsp;Implementation</h1>\n<p>If you have access to a GPU and are using it with a deep learning <a href=\"https://pytorch.org/\" target=\"_blank\">PyTorch</a> model, the best library to use is probably <code>nnAudio</code> since it gives a huge execution time boost since it runs on GPU. The code follows the algorithm proposed in the following <a href=\"http://academics.wellesley.edu/Physics/brown/pubs/effalgV92P2698-P2701.pdf\" target=\"_blank\">paper</a> with a slight improvement.</p>\n<p>The main function that we will be using is <code>CQT1992v2</code>: <strong>1992</strong> stands for the year the algorithm's paper has been published and <strong>v2</strong> stands for a slight improvement on the original paper. </p>\n<p>Here is a code snippet to do it (notice that  the code below will run on CPU): </p>\n<pre><code>import torch\nfrom nnAudio.Spectrogram import CQT1992v2\nimport matplotlib.pylab as plt\nimport numpy as np\n\n# sr is the sampling rate, it is 2048 Hz\n# fmax is half the sampling rate\n\ncqt_transform = CQT1992v2(sr=2048, fmin=20, fmax=1024, hop_length=64)\n\ndef run_cqt_transform(x: np.array) -&gt; torch.Tensor:\n    # We stack the passed x since there are 3\n    # time series per file.\n    x = np.hstack(x)\n    # Normalize (is there a better way?)\n    x = x / np.max(x)\n    x = torch.from_numpy(x).float()\n    return cqt_transform(x)\n\n# Running on one file and plotting the result.\nx = np.load(\"path/to/file.npy\")\n\n# We take the first (and only) result since the \n# result is batch-shaped ((1, freq_bins, time_steps)).\nimg = run_cqt_transform(x)[0]\n\nfig, ax = plt.subplots(1, 1, figsize=(12, 8))\nax.imshow(img)\nplt.show()\n</code></pre>\n<p>Here is a sample of an output that you can get:</p>\n<p><a href=\"https://ibb.co/tDHZ5Ph\"><img src=\"https://i.ibb.co/5L1YVx9/cqt-sample-output.png\" alt=\"cqt-sample-output\"></a></p>\n<p>Notice that there are many other parameters that you can pass to <code>CQT1992v2</code>. <br>\nOne of these is the <code>bins_per_octave</code> which defaults to <code>12</code> and the <code>output_format</code> which defaults to <code>Magnitude</code> (i.e. the complex module). You can also get the real and imaginary parts with <code>Complex</code>.</p>\n<h1>Link to the Morlet wavelet transform</h1>\n<p>If we use the <a href=\"https://en.wikipedia.org/wiki/Morlet_wavelet\" target=\"_blank\">Morlet wavelet</a>, i.e. a windowed (complex) sinusoidal where the window is a Gaussian, in a continuous wavelet transform it becomes a kind of STFT with the Morelet wavelet replacing the window part X[n-m] times the (complex) sinusoidal. Then, doing the same steps as in Theoretical computation section, we get<br>\nan analogous constant-q transform (the filters have the correct scaling thanks to the wavelet's scaling). </p>\n<p>The advantage of the Morlet wavelet transform in comparison with CQT is that its<br>\ninverse is easier to compute since the wavelets form a basis, whereas in CQT it isn't the case in general.</p>\n<p>More details can be found <a href=\"https://ccrma.stanford.edu/~jos/sasp/Continuous_Wavelet_Transform.html\" target=\"_blank\">here</a>.</p>\n<h1>Resources</h1>\n<p>For more details, check the Wikipedia <a href=\"https://en.wikipedia.org/wiki/Constant-Q_transform\" target=\"_blank\">page</a>.</p>\n<p>And finally, for those that like to check even more things, here are links to related concepts and further explanations:</p>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Fourier_transform\" target=\"_blank\">Fourier transform</a></li>\n<li><a href=\"https://en.wikipedia.org/wiki/Short-time_Fourier_transform\" target=\"_blank\">Short-time Fourier transform</a></li>\n<li><a href=\"https://en.wikipedia.org/wiki/Gabor_transform\" target=\"_blank\">Gabor transform</a></li>\n<li><a href=\"https://en.wikipedia.org/wiki/List_of_Fourier-related_transforms\" target=\"_blank\">Many more Fourier related transforms</a></li>\n<li>A <a href=\"https://www.youtube.com/watch?v=EfWnEldTyPA\" target=\"_blank\">video</a> about the spectrogram and the Gabor transform </li>\n<li>A <a href=\"https://www.youtube.com/watch?v=Cl-m4X3rwac\" target=\"_blank\">video</a> about the constant-q transform</li>\n</ul>\n<p>[UPDATE] Few small code fixes thanks to <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> comment. </p>",
      "rawMarkdown": "# Introduction\n\nI have discovered this transformation thanks to the following [notebook](https://www.kaggle.com/atamazian/nnaudio-constant-q-transform-demonstration) by [Araik Tamazian](https://www.kaggle.com/atamazian). \n\nIt is a **signal processing** transformation closely related to the [**Fourier transform**](https://en.wikipedia.org/wiki/Fourier_transform) (more specifically to the [short term](https://en.wikipedia.org/wiki/Short-time_Fourier_transform) version of it). \n\nIt is also related to the complex [Morlet wavelet](https://en.wikipedia.org/wiki/Morlet_wavelet). We will see how in a later section.\n\nAlso, notice that the full name is **constant-Q transform** (or CQT in short). This name comes from the fact\nthat the transformation can be thought of as applying different filters equally spaced on a logarithmic frequency \nscale. The k-th filter is related to the (k-1)-th filter using the following formula (from Wikipeda):\n\n<a href=\"https://ibb.co/Wv4nwdf\"><img src=\"https://i.ibb.co/0yRqSWr/delta-fk-cqt.png\" alt=\"delta-fk-cqt\" border=\"0\"></a>\n\nwhere ''δf''_'k'' is the bandwidth of the ''k''-th filter, ''f''_min is the central frequency of the lowest filter, and ''n'' is the number of filters per octave.\n\n\nThe **CQT** has been used a lot in music since it is well suited to the limited range of human's auditory perception \n(20 to 20k Hz) and to the fact that the human ear is more sensitive to lower frequencies than higher ones.\n\n# Theoretical computation\n\nHere is how it is defined and computed in few steps:\n\n1. We start by computing the [short-term Fourier transform](https://en.wikipedia.org/wiki/Short-time_Fourier_transform) (STFT in short): this is a Fourier transform with a moving window over the signal.\nAs a window function, we often use the [Hann or Hamming windows](https://en.wikipedia.org/wiki/Window_function#Hann_and_Hamming_windows).\nFinally, notice that we will use the discrete version of the STFT. The equation should be: \n\n<a href=\"https://ibb.co/KwNjb7y\"><img src=\"https://i.ibb.co/HdNDBYg/stft-cqt.png\" alt=\"stft-cqt\" border=\"0\"></a>\n\n2. We introduce the **quality factor** `Q` using the following formula: \n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/12ZckR8/q-cqt.png\" alt=\"q-cqt\" border=\"0\"></a>\n\n3. Now, the window length for the k-th bin is no longer fixed (N) by will vary depending on the filter \n(we use the previous definition to get a simplified expression): \n\n<a href=\"https://ibb.co/crwqYgc\"><img src=\"https://i.ibb.co/YLXg0R7/n-cqt.png\" alt=\"n-cqt\" border=\"0\"></a>\n\n3. The digital frequency \\frac{2 \\pi k}{N} is also changed and the window function now depends on k as well (via the window length N[k])\n\n4. Finally, replacing the different elements in the original formula, we get: \n\n\n<a href=\"https://ibb.co/T48ZfYL\"><img src=\"https://i.ibb.co/wSWTkBs/cqt-cqt.png\" alt=\"cqt-cqt\" border=\"0\"></a>\n\n\nNotice that for a fast implementation, [FFT](https://en.wikipedia.org/wiki/Fast_Fourier_transform) is often used. Sometimes, it can even be replaced \nwith a [sliding DFT](https://en.wikipedia.org/wiki/Sliding_DFT).\n\n\n# Implementation\n\nIf you have access to a GPU and are using it with a deep learning [PyTorch](https://pytorch.org/) model, the best library to use is probably `nnAudio` since it gives a huge execution time boost since it runs on GPU. The code follows the algorithm proposed in the following [paper](http://academics.wellesley.edu/Physics/brown/pubs/effalgV92P2698-P2701.pdf) with a slight improvement.\n\nThe main function that we will be using is `CQT1992v2`: **1992** stands for the year the algorithm's paper has been published and **v2** stands for a slight improvement on the original paper. \n\nHere is a code snippet to do it (notice that  the code below will run on CPU): \n\n\n```python\n\n\nimport torch\nfrom nnAudio.Spectrogram import CQT1992v2\nimport matplotlib.pylab as plt\nimport numpy as np\n\n# sr is the sampling rate, it is 2048 Hz\n# fmax is half the sampling rate\n\ncqt_transform = CQT1992v2(sr=2048, fmin=20, fmax=1024, hop_length=64)\n\ndef run_cqt_transform(x: np.array) -> torch.Tensor:\n    # We stack the passed x since there are 3\n    # time series per file.\n    x = np.hstack(x)\n    # Normalize (is there a better way?)\n    x = x / np.max(x)\n    x = torch.from_numpy(x).float()\n    return cqt_transform(x)\n\n# Running on one file and plotting the result.\nx = np.load(\"path/to/file.npy\")\n\n# We take the first (and only) result since the \n# result is batch-shaped ((1, freq_bins, time_steps)).\nimg = run_cqt_transform(x)[0]\n\nfig, ax = plt.subplots(1, 1, figsize=(12, 8))\nax.imshow(img)\nplt.show()\n\n\n```\n\nHere is a sample of an output that you can get:\n\n<a href=\"https://ibb.co/tDHZ5Ph\"><img src=\"https://i.ibb.co/5L1YVx9/cqt-sample-output.png\" alt=\"cqt-sample-output\" border=\"0\"></a>\n\n\nNotice that there are many other parameters that you can pass to `CQT1992v2`. \nOne of these is the `bins_per_octave` which defaults to `12` and the `output_format` which defaults to `Magnitude` (i.e. the complex module). You can also get the real and imaginary parts with `Complex`.\n\n\n# Link to the Morlet wavelet transform\n\nIf we use the [Morlet wavelet](https://en.wikipedia.org/wiki/Morlet_wavelet), i.e. a windowed (complex) sinusoidal where the window is a Gaussian, in a continuous wavelet transform it becomes a kind of STFT with the Morelet wavelet replacing the window part X[n-m] times the (complex) sinusoidal. Then, doing the same steps as in Theoretical computation section, we get\nan analogous constant-q transform (the filters have the correct scaling thanks to the wavelet's scaling). \n\nThe advantage of the Morlet wavelet transform in comparison with CQT is that its\ninverse is easier to compute since the wavelets form a basis, whereas in CQT it isn't the case in general.\n\nMore details can be found [here](https://ccrma.stanford.edu/~jos/sasp/Continuous_Wavelet_Transform.html).\n\n# Resources\n\nFor more details, check the Wikipedia [page](https://en.wikipedia.org/wiki/Constant-Q_transform).\n\n\nAnd finally, for those that like to check even more things, here are links to related concepts and further explanations:\n\n- [Fourier transform](https://en.wikipedia.org/wiki/Fourier_transform)\n- [Short-time Fourier transform](https://en.wikipedia.org/wiki/Short-time_Fourier_transform)\n- [Gabor transform](https://en.wikipedia.org/wiki/Gabor_transform)\n- [Many more Fourier related transforms](https://en.wikipedia.org/wiki/List_of_Fourier-related_transforms)\n- A [video](https://www.youtube.com/watch?v=EfWnEldTyPA) about the spectrogram and the Gabor transform \n- A [video](https://www.youtube.com/watch?v=Cl-m4X3rwac) about the constant-q transform\n\n\n[UPDATE] Few small code fixes thanks to @cpmpml comment.",
      "votes": null
    },
    {
      "id": "1494589",
      "postDate": "08/28/2021 20:00:58",
      "content": "<p>By the way, is there a better way to render Latex equations in Kaggle? I have tried the $ marker but it doesn't seem to work? So far, the best option is to use this website: <a href=\"http://latex2png.com/\" target=\"_blank\">http://latex2png.com/</a>. </p>\n<p>Maybe this is a better option for the rendering inline: <a href=\"https://latex.codecogs.com\" target=\"_blank\">https://latex.codecogs.com</a></p>",
      "rawMarkdown": "By the way, is there a better way to render Latex equations in Kaggle? I have tried the $ marker but it doesn't seem to work? So far, the best option is to use this website: http://latex2png.com/. \n\nMaybe this is a better option for the rendering inline: https://latex.codecogs.com",
      "votes": null
    },
    {
      "id": "1494614",
      "postDate": "08/28/2021 20:56:13",
      "content": "<p>Hi, small mistake but can lead to bad performance. This does not move data to GPU:</p>\n<pre><code>    # Move from CPU to GPU\n    x = torch.from_numpy(x).float()\n</code></pre>",
      "rawMarkdown": "Hi, small mistake but can lead to bad performance. This does not move data to GPU:\n\n```\n    # Move from CPU to GPU\n    x = torch.from_numpy(x).float()\n\n```",
      "votes": null
    },
    {
      "id": "1494618",
      "postDate": "08/28/2021 21:05:31",
      "content": "<p>Thanks for catching this. I will fix it. 👌 </p>",
      "rawMarkdown": "Thanks for catching this. I will fix it. 👌",
      "votes": null
    },
    {
      "id": "1494626",
      "postDate": "08/28/2021 21:28:14",
      "content": "<p>For those that wonder why it isn't correct, by default, when creating a Tensor, it will be on the CPU even if a GPU is available (as far as my understanding goes).</p>\n<p>To have it on the GPU, you will need to either use the <code>.to</code> method or specify the <code>device</code> argument when creating the <code>Tensor</code>.</p>\n<p>Once you have your torch.Tensor object, you can use the property <code>device</code> or the method <code>is_cuda</code> to check where it is located. </p>\n<p>More details can be found in the official documentation: <a href=\"https://pytorch.org/docs/stable/tensors.html\" target=\"_blank\">https://pytorch.org/docs/stable/tensors.html</a></p>",
      "rawMarkdown": "For those that wonder why it isn't correct, by default, when creating a Tensor, it will be on the CPU even if a GPU is available (as far as my understanding goes).\n\nTo have it on the GPU, you will need to either use the `.to` method or specify the `device` argument when creating the `Tensor`.\n\nOnce you have your torch.Tensor object, you can use the property `device` or the method `is_cuda` to check where it is located. \n\nMore details can be found in the official documentation: https://pytorch.org/docs/stable/tensors.html",
      "votes": null
    },
    {
      "id": "1494705",
      "postDate": "08/29/2021 02:17:45",
      "content": "<p>Hi, how is output shape from CQT computed ?</p>\n<p>So in your case the input to CQT will be <code>[N, 4096*3]</code> . And according to the docs the output shape should be <code>[N, freq_bins, time_steps]</code>. How is <code>freq_bins</code> &amp; <code>time_steps</code> computed ?</p>",
      "rawMarkdown": "Hi, how is output shape from CQT computed ?\n\nSo in your case the input to CQT will be `[N, 4096*3]` . And according to the docs the output shape should be `[N, freq_bins, time_steps]`. How is `freq_bins` & `time_steps` computed ?",
      "votes": null
    },
    {
      "id": "1494948",
      "postDate": "08/29/2021 07:33:43",
      "content": "<p>Not sure that I understand your question but I will try to  provide some details, hopefully it helps.</p>\n<p>If you check this paper and the <a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/Spectrogram.py\" target=\"_blank\">source code</a>, you will notice that FFT coefficients are computed first then a complex multiplication of these FFT coefficients is done since this makes the CQT computation faster </p>\n<p>This is equivalent to the original equation thanks to <a href=\"https://en.wikipedia.org/wiki/Parseval%27s_identity\" target=\"_blank\">Parseval</a>'s identity. </p>\n<p>Notice that this is the original CQT1992. The CQT1992v2 uses a slightly different computation (using <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Conv1d.html\" target=\"_blank\">conv1D</a> instead). </p>\n<p>As for the frequency bins and time steps, these are derived using the <code>create_cqt_kernels function</code> (notice that this function is used on both implementations of CQT). You can find the source code <a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/0c4f77746839d3db4fbabbfd3d51a897c5845856/Installation/nnAudio/utils.py#L309\" target=\"_blank\">here</a>.</p>\n<p>Let me know if this helps a bit. 😀</p>",
      "rawMarkdown": "Not sure that I understand your question but I will try to  provide some details, hopefully it helps.\n\nIf you check this paper and the [source code](https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/Spectrogram.py), you will notice that FFT coefficients are computed first then a complex multiplication of these FFT coefficients is done since this makes the CQT computation faster \n\nThis is equivalent to the original equation thanks to [Parseval](https://en.wikipedia.org/wiki/Parseval%27s_identity)'s identity. \n\nNotice that this is the original CQT1992. The CQT1992v2 uses a slightly different computation (using [conv1D](https://pytorch.org/docs/stable/generated/torch.nn.Conv1d.html) instead). \n\nAs for the frequency bins and time steps, these are derived using the `create_cqt_kernels function` (notice that this function is used on both implementations of CQT). You can find the source code [here](https://github.com/KinWaiCheuk/nnAudio/blob/0c4f77746839d3db4fbabbfd3d51a897c5845856/Installation/nnAudio/utils.py#L309).\n\nLet me know if this helps a bit. 😀",
      "votes": null
    },
    {
      "id": "1494954",
      "postDate": "08/29/2021 07:40:12",
      "content": "<p><code>time_steps</code> is controlled by the signal length (4096) / <code>hop_length</code></p>\n<p><code>freq_bins</code> is controlled by <code>bins_per_octave</code> and <code>filter_scale</code> (see the docs for the <code>filter_scale</code> arg to see how these two are related). But if you are using <code>fmin</code> and <code>fmax</code>, some of the bins are removed.</p>",
      "rawMarkdown": "`time_steps` is controlled by the signal length (4096) / `hop_length`\n\n`freq_bins` is controlled by `bins_per_octave` and `filter_scale` (see the docs for the `filter_scale` arg to see how these two are related). But if you are using `fmin` and `fmax`, some of the bins are removed.",
      "votes": null
    },
    {
      "id": "1494958",
      "postDate": "08/29/2021 07:44:17",
      "content": "<p>I think that your answer is better <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, I think I finally understood what <a href=\"https://www.kaggle.com/benihime91\" target=\"_blank\">@benihime91</a> was asking. 👌</p>",
      "rawMarkdown": "I think that your answer is better @anjum48, I think I finally understood what @benihime91 was asking. 👌",
      "votes": null
    },
    {
      "id": "1495089",
      "postDate": "08/29/2021 09:48:48",
      "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> thanks for the nice explanation .. Been racking my brain around this</p>",
      "rawMarkdown": "anjum48 thanks for the nice explanation .. Been racking my brain around this",
      "votes": null
    },
    {
      "id": "1495879",
      "postDate": "08/29/2021 21:19:04",
      "content": "<p>Very helpful =))</p>",
      "rawMarkdown": "Very helpful =))",
      "votes": null
    },
    {
      "id": "1495887",
      "postDate": "08/29/2021 21:31:25",
      "content": "<p>I am glad it helps!</p>",
      "rawMarkdown": "I am glad it helps!",
      "votes": null
    },
    {
      "id": "1498434",
      "postDate": "09/01/2021 01:40:43",
      "content": "<p>Nice summary. I've been using CWT (Complex Morlet) from using this TF2.0 implementation here <a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">https://github.com/Kevin-McIsaac/cmorlet-tensorflow</a> which runs very well on GPU and TPU.</p>",
      "rawMarkdown": "Nice summary. I've been using CWT (Complex Morlet) from using this TF2.0 implementation here https://github.com/Kevin-McIsaac/cmorlet-tensorflow which runs very well on GPU and TPU.",
      "votes": null
    },
    {
      "id": "1498705",
      "postDate": "09/01/2021 06:23:06",
      "content": "<p>Thanks for sharing this implementation. For PyTorch, I've found this implementation: <a href=\"https://github.com/tomrunia/PyTorchWavelets\" target=\"_blank\">https://github.com/tomrunia/PyTorchWavelets</a>.</p>\n<p>The installation isn't very smooth but there is this notebook that shows how to do it: <a href=\"https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration\" target=\"_blank\">https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration</a></p>\n<p>There are of course many more other implementations: </p>\n<ul>\n<li><a href=\"https://pytorch-wavelets.readthedocs.io/en/latest/readme.html\" target=\"_blank\">https://pytorch-wavelets.readthedocs.io/en/latest/readme.html</a> (PyTorch)</li>\n<li><a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.cwt.html\" target=\"_blank\">https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.cwt.html</a> (CPU) </li>\n</ul>",
      "rawMarkdown": "Thanks for sharing this implementation. For PyTorch, I've found this implementation: https://github.com/tomrunia/PyTorchWavelets.\n\nThe installation isn't very smooth but there is this notebook that shows how to do it: https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration\n\nThere are of course many more other implementations: \n\n- https://pytorch-wavelets.readthedocs.io/en/latest/readme.html (PyTorch)\n- https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.cwt.html (CPU)",
      "votes": null
    },
    {
      "id": "1503749",
      "postDate": "09/05/2021 17:13:12",
      "content": "<p>Basically it's a way of binning the Fourier transformed data.</p>",
      "rawMarkdown": "Basically it's a way of binning the Fourier transformed data.",
      "votes": null
    },
    {
      "id": "1503771",
      "postDate": "09/05/2021 17:32:36",
      "content": "<p>Yes it is. This can be said to a lot of signal processing transformations as well. 😄</p>",
      "rawMarkdown": "Yes it is. This can be said to a lot of signal processing transformations as well. 😄",
      "votes": null
    },
    {
      "id": "1510759",
      "postDate": "09/12/2021 18:06:10",
      "content": "<p>Few things about the <code>hop_length</code> parameter: </p>\n<ul>\n<li>as explained <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270019\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a>, it can be considered as the stride if you think about CQT as convolution (which is the case for the last step). Here is a good illustration shared by him as well:</li>\n</ul>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/MDxfKbf/hop-length.webp\" alt=\"hop-length\"></a></p>\n<ul>\n<li>smaller values of it will give bigger outputs. This could be useful if you want to get finer details.</li>\n</ul>",
      "rawMarkdown": "Few things about the `hop_length` parameter: \n\n- as explained [here](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270019) by @analokamus, it can be considered as the stride if you think about CQT as convolution (which is the case for the last step). Here is a good illustration shared by him as well:\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/MDxfKbf/hop-length.webp\" alt=\"hop-length\" border=\"0\"></a>\n\n- smaller values of it will give bigger outputs. This could be useful if you want to get finer details.",
      "votes": null
    },
    {
      "id": "1510815",
      "postDate": "09/12/2021 18:50:43",
      "content": "<p>Apparently LaTeX can be used<br>\n$$<br>\n\\sum_{i} |f_i|^p \\leq C \\sum_{i,j} |f_i - f_j|^p, \\quad \\text{ if } \\sum_i f_i =0, \\text{ and } p&gt;1 \\text{ fixed}.<br>\n$$<br>\nBoth inline \\(\\LaTeX\\) which uses <code>\\\\( \\\\)</code> and <code>$$ $$</code>.</p>",
      "rawMarkdown": "Apparently LaTeX can be used\n$$\n\\sum_{i} |f_i|^p \\leq C \\sum_{i,j} |f_i - f_j|^p, \\quad \\text{ if } \\sum_i f_i =0, \\text{ and } p>1 \\text{ fixed}.\n$$\nBoth inline \\\\(\\LaTeX\\\\) which uses `\\\\( \\\\)` and `$$ $$`.",
      "votes": null
    },
    {
      "id": "1510838",
      "postDate": "09/12/2021 19:18:36",
      "content": "<p>Oh, so it is double $. I will give it a try, thanks!</p>",
      "rawMarkdown": "Oh, so it is double $. I will give it a try, thanks!",
      "votes": null
    },
    {
      "id": "1559704",
      "postDate": "10/27/2021 07:08:01",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1494589,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "08/28/2021 20:00:58",
      "content": "<p>By the way, is there a better way to render Latex equations in Kaggle? I have tried the $ marker but it doesn't seem to work? So far, the best option is to use this website: <a href=\"http://latex2png.com/\" target=\"_blank\">http://latex2png.com/</a>. </p>\n<p>Maybe this is a better option for the rendering inline: <a href=\"https://latex.codecogs.com\" target=\"_blank\">https://latex.codecogs.com</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1510815,
          "author_name": "scaomath",
          "author_url": "",
          "post_date": "09/12/2021 18:50:43",
          "content": "<p>Apparently LaTeX can be used<br>\n$$<br>\n\\sum_{i} |f_i|^p \\leq C \\sum_{i,j} |f_i - f_j|^p, \\quad \\text{ if } \\sum_i f_i =0, \\text{ and } p&gt;1 \\text{ fixed}.<br>\n$$<br>\nBoth inline \\(\\LaTeX\\) which uses <code>\\\\( \\\\)</code> and <code>$$ $$</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1510838,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/12/2021 19:18:36",
          "content": "<p>Oh, so it is double $. I will give it a try, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1494614,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/28/2021 20:56:13",
      "content": "<p>Hi, small mistake but can lead to bad performance. This does not move data to GPU:</p>\n<pre><code>    # Move from CPU to GPU\n    x = torch.from_numpy(x).float()\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1494618,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/28/2021 21:05:31",
          "content": "<p>Thanks for catching this. I will fix it. 👌 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494626,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/28/2021 21:28:14",
          "content": "<p>For those that wonder why it isn't correct, by default, when creating a Tensor, it will be on the CPU even if a GPU is available (as far as my understanding goes).</p>\n<p>To have it on the GPU, you will need to either use the <code>.to</code> method or specify the <code>device</code> argument when creating the <code>Tensor</code>.</p>\n<p>Once you have your torch.Tensor object, you can use the property <code>device</code> or the method <code>is_cuda</code> to check where it is located. </p>\n<p>More details can be found in the official documentation: <a href=\"https://pytorch.org/docs/stable/tensors.html\" target=\"_blank\">https://pytorch.org/docs/stable/tensors.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1494705,
      "author_name": "benihime91",
      "author_url": "",
      "post_date": "08/29/2021 02:17:45",
      "content": "<p>Hi, how is output shape from CQT computed ?</p>\n<p>So in your case the input to CQT will be <code>[N, 4096*3]</code> . And according to the docs the output shape should be <code>[N, freq_bins, time_steps]</code>. How is <code>freq_bins</code> &amp; <code>time_steps</code> computed ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1494948,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/29/2021 07:33:43",
          "content": "<p>Not sure that I understand your question but I will try to  provide some details, hopefully it helps.</p>\n<p>If you check this paper and the <a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/Spectrogram.py\" target=\"_blank\">source code</a>, you will notice that FFT coefficients are computed first then a complex multiplication of these FFT coefficients is done since this makes the CQT computation faster </p>\n<p>This is equivalent to the original equation thanks to <a href=\"https://en.wikipedia.org/wiki/Parseval%27s_identity\" target=\"_blank\">Parseval</a>'s identity. </p>\n<p>Notice that this is the original CQT1992. The CQT1992v2 uses a slightly different computation (using <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Conv1d.html\" target=\"_blank\">conv1D</a> instead). </p>\n<p>As for the frequency bins and time steps, these are derived using the <code>create_cqt_kernels function</code> (notice that this function is used on both implementations of CQT). You can find the source code <a href=\"https://github.com/KinWaiCheuk/nnAudio/blob/0c4f77746839d3db4fbabbfd3d51a897c5845856/Installation/nnAudio/utils.py#L309\" target=\"_blank\">here</a>.</p>\n<p>Let me know if this helps a bit. 😀</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494954,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "08/29/2021 07:40:12",
          "content": "<p><code>time_steps</code> is controlled by the signal length (4096) / <code>hop_length</code></p>\n<p><code>freq_bins</code> is controlled by <code>bins_per_octave</code> and <code>filter_scale</code> (see the docs for the <code>filter_scale</code> arg to see how these two are related). But if you are using <code>fmin</code> and <code>fmax</code>, some of the bins are removed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494958,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/29/2021 07:44:17",
          "content": "<p>I think that your answer is better <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, I think I finally understood what <a href=\"https://www.kaggle.com/benihime91\" target=\"_blank\">@benihime91</a> was asking. 👌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1495089,
          "author_name": "benihime91",
          "author_url": "",
          "post_date": "08/29/2021 09:48:48",
          "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> thanks for the nice explanation .. Been racking my brain around this</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1495879,
      "author_name": "top10chi3nthan",
      "author_url": "",
      "post_date": "08/29/2021 21:19:04",
      "content": "<p>Very helpful =))</p>",
      "votes": null,
      "replies": [
        {
          "id": 1495887,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/29/2021 21:31:25",
          "content": "<p>I am glad it helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1498434,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "09/01/2021 01:40:43",
      "content": "<p>Nice summary. I've been using CWT (Complex Morlet) from using this TF2.0 implementation here <a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">https://github.com/Kevin-McIsaac/cmorlet-tensorflow</a> which runs very well on GPU and TPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1498705,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/01/2021 06:23:06",
          "content": "<p>Thanks for sharing this implementation. For PyTorch, I've found this implementation: <a href=\"https://github.com/tomrunia/PyTorchWavelets\" target=\"_blank\">https://github.com/tomrunia/PyTorchWavelets</a>.</p>\n<p>The installation isn't very smooth but there is this notebook that shows how to do it: <a href=\"https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration\" target=\"_blank\">https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration</a></p>\n<p>There are of course many more other implementations: </p>\n<ul>\n<li><a href=\"https://pytorch-wavelets.readthedocs.io/en/latest/readme.html\" target=\"_blank\">https://pytorch-wavelets.readthedocs.io/en/latest/readme.html</a> (PyTorch)</li>\n<li><a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.cwt.html\" target=\"_blank\">https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.cwt.html</a> (CPU) </li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1503749,
      "author_name": "tolgadincer",
      "author_url": "",
      "post_date": "09/05/2021 17:13:12",
      "content": "<p>Basically it's a way of binning the Fourier transformed data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1503771,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "09/05/2021 17:32:36",
          "content": "<p>Yes it is. This can be said to a lot of signal processing transformations as well. 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1510759,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "09/12/2021 18:06:10",
      "content": "<p>Few things about the <code>hop_length</code> parameter: </p>\n<ul>\n<li>as explained <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270019\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a>, it can be considered as the stride if you think about CQT as convolution (which is the case for the last step). Here is a good illustration shared by him as well:</li>\n</ul>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/MDxfKbf/hop-length.webp\" alt=\"hop-length\"></a></p>\n<ul>\n<li>smaller values of it will give bigger outputs. This could be useful if you want to get finer details.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559704,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:08:01",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1494588": "# Introduction\n\nI have discovered this transformation thanks to the following [notebook](https://www.kaggle.com/atamazian/nnaudio-constant-q-transform-demonstration) by [Araik Tamazian](https://www.kaggle.com/atamazian). \n\nIt is a **signal processing** transformation closely related to the [**Fourier transform**](https://en.wikipedia.org/wiki/Fourier_transform) (more specifically to the [short term](https://en.wikipedia.org/wiki/Short-time_Fourier_transform) version of it). \n\nIt is also related to the complex [Morlet wavelet](https://en.wikipedia.org/wiki/Morlet_wavelet). We will see how in a later section.\n\nAlso, notice that the full name is **constant-Q transform** (or CQT in short). This name comes from the fact\nthat the transformation can be thought of as applying different filters equally spaced on a logarithmic frequency \nscale. The k-th filter is related to the (k-1)-th filter using the following formula (from Wikipeda):\n\n<a href=\"https://ibb.co/Wv4nwdf\"><img src=\"https://i.ibb.co/0yRqSWr/delta-fk-cqt.png\" alt=\"delta-fk-cqt\" border=\"0\"></a>\n\nwhere ''δf''_'k'' is the bandwidth of the ''k''-th filter, ''f''_min is the central frequency of the lowest filter, and ''n'' is the number of filters per octave.\n\n\nThe **CQT** has been used a lot in music since it is well suited to the limited range of human's auditory perception \n(20 to 20k Hz) and to the fact that the human ear is more sensitive to lower frequencies than higher ones.\n\n# Theoretical computation\n\nHere is how it is defined and computed in few steps:\n\n1. We start by computing the [short-term Fourier transform](https://en.wikipedia.org/wiki/Short-time_Fourier_transform) (STFT in short): this is a Fourier transform with a moving window over the signal.\nAs a window function, we often use the [Hann or Hamming windows](https://en.wikipedia.org/wiki/Window_function#Hann_and_Hamming_windows).\nFinally, notice that we will use the discrete version of the STFT. The equation should be: \n\n<a href=\"https://ibb.co/KwNjb7y\"><img src=\"https://i.ibb.co/HdNDBYg/stft-cqt.png\" alt=\"stft-cqt\" border=\"0\"></a>\n\n2. We introduce the **quality factor** `Q` using the following formula: \n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/12ZckR8/q-cqt.png\" alt=\"q-cqt\" border=\"0\"></a>\n\n3. Now, the window length for the k-th bin is no longer fixed (N) by will vary depending on the filter \n(we use the previous definition to get a simplified expression): \n\n<a href=\"https://ibb.co/crwqYgc\"><img src=\"https://i.ibb.co/YLXg0R7/n-cqt.png\" alt=\"n-cqt\" border=\"0\"></a>\n\n3. The digital frequency \\frac{2 \\pi k}{N} is also changed and the window function now depends on k as well (via the window length N[k])\n\n4. Finally, replacing the different elements in the original formula, we get: \n\n\n<a href=\"https://ibb.co/T48ZfYL\"><img src=\"https://i.ibb.co/wSWTkBs/cqt-cqt.png\" alt=\"cqt-cqt\" border=\"0\"></a>\n\n\nNotice that for a fast implementation, [FFT](https://en.wikipedia.org/wiki/Fast_Fourier_transform) is often used. Sometimes, it can even be replaced \nwith a [sliding DFT](https://en.wikipedia.org/wiki/Sliding_DFT).\n\n\n# Implementation\n\nIf you have access to a GPU and are using it with a deep learning [PyTorch](https://pytorch.org/) model, the best library to use is probably `nnAudio` since it gives a huge execution time boost since it runs on GPU. The code follows the algorithm proposed in the following [paper](http://academics.wellesley.edu/Physics/brown/pubs/effalgV92P2698-P2701.pdf) with a slight improvement.\n\nThe main function that we will be using is `CQT1992v2`: **1992** stands for the year the algorithm's paper has been published and **v2** stands for a slight improvement on the original paper. \n\nHere is a code snippet to do it (notice that  the code below will run on CPU): \n\n\n```python\n\n\nimport torch\nfrom nnAudio.Spectrogram import CQT1992v2\nimport matplotlib.pylab as plt\nimport numpy as np\n\n# sr is the sampling rate, it is 2048 Hz\n# fmax is half the sampling rate\n\ncqt_transform = CQT1992v2(sr=2048, fmin=20, fmax=1024, hop_length=64)\n\ndef run_cqt_transform(x: np.array) -> torch.Tensor:\n    # We stack the passed x since there are 3\n    # time series per file.\n    x = np.hstack(x)\n    # Normalize (is there a better way?)\n    x = x / np.max(x)\n    x = torch.from_numpy(x).float()\n    return cqt_transform(x)\n\n# Running on one file and plotting the result.\nx = np.load(\"path/to/file.npy\")\n\n# We take the first (and only) result since the \n# result is batch-shaped ((1, freq_bins, time_steps)).\nimg = run_cqt_transform(x)[0]\n\nfig, ax = plt.subplots(1, 1, figsize=(12, 8))\nax.imshow(img)\nplt.show()\n\n\n```\n\nHere is a sample of an output that you can get:\n\n<a href=\"https://ibb.co/tDHZ5Ph\"><img src=\"https://i.ibb.co/5L1YVx9/cqt-sample-output.png\" alt=\"cqt-sample-output\" border=\"0\"></a>\n\n\nNotice that there are many other parameters that you can pass to `CQT1992v2`. \nOne of these is the `bins_per_octave` which defaults to `12` and the `output_format` which defaults to `Magnitude` (i.e. the complex module). You can also get the real and imaginary parts with `Complex`.\n\n\n# Link to the Morlet wavelet transform\n\nIf we use the [Morlet wavelet](https://en.wikipedia.org/wiki/Morlet_wavelet), i.e. a windowed (complex) sinusoidal where the window is a Gaussian, in a continuous wavelet transform it becomes a kind of STFT with the Morelet wavelet replacing the window part X[n-m] times the (complex) sinusoidal. Then, doing the same steps as in Theoretical computation section, we get\nan analogous constant-q transform (the filters have the correct scaling thanks to the wavelet's scaling). \n\nThe advantage of the Morlet wavelet transform in comparison with CQT is that its\ninverse is easier to compute since the wavelets form a basis, whereas in CQT it isn't the case in general.\n\nMore details can be found [here](https://ccrma.stanford.edu/~jos/sasp/Continuous_Wavelet_Transform.html).\n\n# Resources\n\nFor more details, check the Wikipedia [page](https://en.wikipedia.org/wiki/Constant-Q_transform).\n\n\nAnd finally, for those that like to check even more things, here are links to related concepts and further explanations:\n\n- [Fourier transform](https://en.wikipedia.org/wiki/Fourier_transform)\n- [Short-time Fourier transform](https://en.wikipedia.org/wiki/Short-time_Fourier_transform)\n- [Gabor transform](https://en.wikipedia.org/wiki/Gabor_transform)\n- [Many more Fourier related transforms](https://en.wikipedia.org/wiki/List_of_Fourier-related_transforms)\n- A [video](https://www.youtube.com/watch?v=EfWnEldTyPA) about the spectrogram and the Gabor transform \n- A [video](https://www.youtube.com/watch?v=Cl-m4X3rwac) about the constant-q transform\n\n\n[UPDATE] Few small code fixes thanks to @cpmpml comment.",
    "1494589": "By the way, is there a better way to render Latex equations in Kaggle? I have tried the $ marker but it doesn't seem to work? So far, the best option is to use this website: http://latex2png.com/. \n\nMaybe this is a better option for the rendering inline: https://latex.codecogs.com",
    "1494614": "Hi, small mistake but can lead to bad performance. This does not move data to GPU:\n\n```\n    # Move from CPU to GPU\n    x = torch.from_numpy(x).float()\n\n```",
    "1494618": "Thanks for catching this. I will fix it. 👌",
    "1494626": "For those that wonder why it isn't correct, by default, when creating a Tensor, it will be on the CPU even if a GPU is available (as far as my understanding goes).\n\nTo have it on the GPU, you will need to either use the `.to` method or specify the `device` argument when creating the `Tensor`.\n\nOnce you have your torch.Tensor object, you can use the property `device` or the method `is_cuda` to check where it is located. \n\nMore details can be found in the official documentation: https://pytorch.org/docs/stable/tensors.html",
    "1494705": "Hi, how is output shape from CQT computed ?\n\nSo in your case the input to CQT will be `[N, 4096*3]` . And according to the docs the output shape should be `[N, freq_bins, time_steps]`. How is `freq_bins` & `time_steps` computed ?",
    "1494948": "Not sure that I understand your question but I will try to  provide some details, hopefully it helps.\n\nIf you check this paper and the [source code](https://github.com/KinWaiCheuk/nnAudio/blob/master/Installation/nnAudio/Spectrogram.py), you will notice that FFT coefficients are computed first then a complex multiplication of these FFT coefficients is done since this makes the CQT computation faster \n\nThis is equivalent to the original equation thanks to [Parseval](https://en.wikipedia.org/wiki/Parseval%27s_identity)'s identity. \n\nNotice that this is the original CQT1992. The CQT1992v2 uses a slightly different computation (using [conv1D](https://pytorch.org/docs/stable/generated/torch.nn.Conv1d.html) instead). \n\nAs for the frequency bins and time steps, these are derived using the `create_cqt_kernels function` (notice that this function is used on both implementations of CQT). You can find the source code [here](https://github.com/KinWaiCheuk/nnAudio/blob/0c4f77746839d3db4fbabbfd3d51a897c5845856/Installation/nnAudio/utils.py#L309).\n\nLet me know if this helps a bit. 😀",
    "1494954": "`time_steps` is controlled by the signal length (4096) / `hop_length`\n\n`freq_bins` is controlled by `bins_per_octave` and `filter_scale` (see the docs for the `filter_scale` arg to see how these two are related). But if you are using `fmin` and `fmax`, some of the bins are removed.",
    "1494958": "I think that your answer is better @anjum48, I think I finally understood what @benihime91 was asking. 👌",
    "1495089": "anjum48 thanks for the nice explanation .. Been racking my brain around this",
    "1495879": "Very helpful =))",
    "1495887": "I am glad it helps!",
    "1498434": "Nice summary. I've been using CWT (Complex Morlet) from using this TF2.0 implementation here https://github.com/Kevin-McIsaac/cmorlet-tensorflow which runs very well on GPU and TPU.",
    "1498705": "Thanks for sharing this implementation. For PyTorch, I've found this implementation: https://github.com/tomrunia/PyTorchWavelets.\n\nThe installation isn't very smooth but there is this notebook that shows how to do it: https://www.kaggle.com/atamazian/pytorchwavelets-cwt-demonstration\n\nThere are of course many more other implementations: \n\n- https://pytorch-wavelets.readthedocs.io/en/latest/readme.html (PyTorch)\n- https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.cwt.html (CPU)",
    "1503749": "Basically it's a way of binning the Fourier transformed data.",
    "1503771": "Yes it is. This can be said to a lot of signal processing transformations as well. 😄",
    "1510759": "Few things about the `hop_length` parameter: \n\n- as explained [here](https://www.kaggle.com/c/g2net-gravitational-wave-detection/discussion/270019) by @analokamus, it can be considered as the stride if you think about CQT as convolution (which is the case for the last step). Here is a good illustration shared by him as well:\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/MDxfKbf/hop-length.webp\" alt=\"hop-length\" border=\"0\"></a>\n\n- smaller values of it will give bigger outputs. This could be useful if you want to get finer details.",
    "1510815": "Apparently LaTeX can be used\n$$\n\\sum_{i} |f_i|^p \\leq C \\sum_{i,j} |f_i - f_j|^p, \\quad \\text{ if } \\sum_i f_i =0, \\text{ and } p>1 \\text{ fixed}.\n$$\nBoth inline \\\\(\\LaTeX\\\\) which uses `\\\\( \\\\)` and `$$ $$`.",
    "1510838": "Oh, so it is double $. I will give it a try, thanks!",
    "1559704": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}