{
  "id": 487110,
  "title": "Sorry Rapids, not this time",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/487110",
  "author_name": "SSS",
  "post_date": "2024-03-27T16:41:11.275000",
  "votes": 52,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<p>First things first, I'd like to say thank you to <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> and his team for data preparation idea posted <a href=\"https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\" target=\"_blank\">in this notebook</a>. Honestly, it is how we get our current <strong>0.28</strong> on the leaderboard. Though, I'd like to offer some improvements to the main script function <code>create_spectrogram_with_cusignal</code>. </p>\n<h2>The script</h2>\n<p>Rafael creates spectrogram by using Rapids, namely cuSignal, which might be a bit painful to install and keepup with versions. The exact same code can be rewritten  by using preinstalled cupy library (hasstle free solution). Here is the whole thing:</p>\n<pre><code>\n\n\n\n\n cupy  cp\n numpy  np\n pandas  pd\n cupyx.scipy.ndimage  gaussian_filter\n cupyx.scipy.signal  filtfilt, iirnotch\n cupyx.scipy.signal  spectrogram  cupyx_spectrogram\n scipy.signal  filtfilt  scipy_filtfilt, butter  scipy_butter\n\n ():\n    electrode_pair_name_locations = {: [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ]}\n\n    \n    nyquist_freq =  * \n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    \n    \n    notch_coefficients = iirnotch(w0=, Q=, fs=)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype=)\n\n    spec_size = duration * \n    start = start * \n    real_start = start + ( // ) - (spec_size // )\n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    \n    fs = \n\n     spec_size_freq &lt;=   spec_size_time &lt;= :\n        freq_size = ((nfft // ) / ) + \n        segments = ((spec_size - noverlap) / (nperseg - noverlap))\n    :\n        freq_size = spec_size_freq\n        segments = spec_size_time\n\n    \n    spectrogram = cp.zeros((freq_size, segments, ), dtype=)\n\n    processed_eeg = {}\n\n     i, (electrode_pair_name, electrode_locs)  (electrode_pair_name_locations.items()):\n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n         j  ():\n            \n            signal = cp.array(eeg_data[electrode_locs[j]].values - eeg_data[electrode_locs[j + ]].values)\n\n            \n            mean_signal = cp.nanmean(signal)\n            signal = cp.nan_to_num(signal, nan=mean_signal)  cp.isnan(signal).mean() &lt;   cp.zeros_like(signal)\n\n            \n            signal_filtered = filtfilt(*notch_coefficients, signal)\n            signal_filtered = scipy_filtfilt(*sci_bandpass_coefficients, signal_filtered.get())  \n\n            \n            frequencies, times, Sxx = cupyx_spectrogram(signal_filtered, fs, nperseg=nperseg, noverlap=noverlap,\n                                                        nfft=nfft)\n\n            \n            valid_freq = (frequencies &gt;= ) &amp; (frequencies &lt;= )\n            Sxx_filtered = Sxx[valid_freq, :]\n\n            \n            spectrogram_slice = cp.clip(Sxx_filtered, cp.exp(-), cp.exp())\n            spectrogram_slice = cp.log10(spectrogram_slice)\n\n            normalization_epsilon = \n            mean = spectrogram_slice.mean(axis=(, ), keepdims=)\n            std = spectrogram_slice.std(axis=(, ), keepdims=)\n            spectrogram_slice = (spectrogram_slice - mean) / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        \n         mean_montage_names &gt; :\n            spectrogram[:, :, i] /= mean_montage_names\n\n    \n    spec_numpy = gaussian_filter(spectrogram, sigma=sigma_gaussian).get()  sigma_gaussian &gt;   spectrogram.get()\n\n    \n    ekg_signal_filtered = filtfilt(*notch_coefficients, cp.array(eeg_data[].values))\n    processed_eeg[] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered.get())  \n     spec_numpy, processed_eeg\n</code></pre>\n<h2>Bonus</h2>\n<p>Also, after preparing 50 sec and 10 sec specs you can create the final image differently. So for example, rather than resizing it you can cut high frequencies like shown below and center left-right parts mirror-like.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Feefd98c6ccc8b8a1afb782d0952175fe%2F1.png?generation=1711557383664757&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ffeb52985983b67376f4aeef64742937e%2F2.png?generation=1711557395408198&amp;alt=media\"></p>\n<p>Though the originally proposed method used by Rafael brings the best score for us. You are welcome to share your ideas on how you combine the thing, but I am pretty much sure top teams did not leave Rafael's notebook unnoticed.</p>\n<h2>Outro</h2>\n<p>Have fun and good luck, there is still a time to improve!</p>",
  "messages": [
    {
      "id": 2719349,
      "postDate": "2024-03-27T16:41:11.277Z",
      "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<p>First things first, I'd like to say thank you to <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a> and his team for data preparation idea posted <a href=\"https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\" target=\"_blank\">in this notebook</a>. Honestly, it is how we get our current <strong>0.28</strong> on the leaderboard. Though, I'd like to offer some improvements to the main script function <code>create_spectrogram_with_cusignal</code>. </p>\n<h2>The script</h2>\n<p>Rafael creates spectrogram by using Rapids, namely cuSignal, which might be a bit painful to install and keepup with versions. The exact same code can be rewritten  by using preinstalled cupy library (hasstle free solution). Here is the whole thing:</p>\n<pre><code>\n\n\n\n\n cupy  cp\n numpy  np\n pandas  pd\n cupyx.scipy.ndimage  gaussian_filter\n cupyx.scipy.signal  filtfilt, iirnotch\n cupyx.scipy.signal  spectrogram  cupyx_spectrogram\n scipy.signal  filtfilt  scipy_filtfilt, butter  scipy_butter\n\n ():\n    electrode_pair_name_locations = {: [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ]}\n\n    \n    nyquist_freq =  * \n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    \n    \n    notch_coefficients = iirnotch(w0=, Q=, fs=)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype=)\n\n    spec_size = duration * \n    start = start * \n    real_start = start + ( // ) - (spec_size // )\n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    \n    fs = \n\n     spec_size_freq &lt;=   spec_size_time &lt;= :\n        freq_size = ((nfft // ) / ) + \n        segments = ((spec_size - noverlap) / (nperseg - noverlap))\n    :\n        freq_size = spec_size_freq\n        segments = spec_size_time\n\n    \n    spectrogram = cp.zeros((freq_size, segments, ), dtype=)\n\n    processed_eeg = {}\n\n     i, (electrode_pair_name, electrode_locs)  (electrode_pair_name_locations.items()):\n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n         j  ():\n            \n            signal = cp.array(eeg_data[electrode_locs[j]].values - eeg_data[electrode_locs[j + ]].values)\n\n            \n            mean_signal = cp.nanmean(signal)\n            signal = cp.nan_to_num(signal, nan=mean_signal)  cp.isnan(signal).mean() &lt;   cp.zeros_like(signal)\n\n            \n            signal_filtered = filtfilt(*notch_coefficients, signal)\n            signal_filtered = scipy_filtfilt(*sci_bandpass_coefficients, signal_filtered.get())  \n\n            \n            frequencies, times, Sxx = cupyx_spectrogram(signal_filtered, fs, nperseg=nperseg, noverlap=noverlap,\n                                                        nfft=nfft)\n\n            \n            valid_freq = (frequencies &gt;= ) &amp; (frequencies &lt;= )\n            Sxx_filtered = Sxx[valid_freq, :]\n\n            \n            spectrogram_slice = cp.clip(Sxx_filtered, cp.exp(-), cp.exp())\n            spectrogram_slice = cp.log10(spectrogram_slice)\n\n            normalization_epsilon = \n            mean = spectrogram_slice.mean(axis=(, ), keepdims=)\n            std = spectrogram_slice.std(axis=(, ), keepdims=)\n            spectrogram_slice = (spectrogram_slice - mean) / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        \n         mean_montage_names &gt; :\n            spectrogram[:, :, i] /= mean_montage_names\n\n    \n    spec_numpy = gaussian_filter(spectrogram, sigma=sigma_gaussian).get()  sigma_gaussian &gt;   spectrogram.get()\n\n    \n    ekg_signal_filtered = filtfilt(*notch_coefficients, cp.array(eeg_data[].values))\n    processed_eeg[] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered.get())  \n     spec_numpy, processed_eeg\n</code></pre>\n<h2>Bonus</h2>\n<p>Also, after preparing 50 sec and 10 sec specs you can create the final image differently. So for example, rather than resizing it you can cut high frequencies like shown below and center left-right parts mirror-like.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Feefd98c6ccc8b8a1afb782d0952175fe%2F1.png?generation=1711557383664757&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ffeb52985983b67376f4aeef64742937e%2F2.png?generation=1711557395408198&amp;alt=media\"></p>\n<p>Though the originally proposed method used by Rafael brings the best score for us. You are welcome to share your ideas on how you combine the thing, but I am pretty much sure top teams did not leave Rafael's notebook unnoticed.</p>\n<h2>Outro</h2>\n<p>Have fun and good luck, there is still a time to improve!</p>",
      "rawMarkdown": "Hi,\n\n##Intro\nFirst things first, I'd like to say thank you to @rafaelzimmermann1 and his team for data preparation idea posted [in this notebook](https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu). Honestly, it is how we get our current **0.28** on the leaderboard. Though, I'd like to offer some improvements to the main script function `create_spectrogram_with_cusignal`. \n\n##The script\nRafael creates spectrogram by using Rapids, namely cuSignal, which might be a bit painful to install and keepup with versions. The exact same code can be rewritten  by using preinstalled cupy library (hasstle free solution). Here is the whole thing:\n\n```python\n# HOTFIX for cupy.signal.filtfilt with bandpass_coefficients cannot find the reason...\n# It produces different results than scipy.signal.filtfilt with the same filter coefficients.\n# signal_filtered = filtfilt(*bandpass_coefficients, signal_filtered)\n\n\nimport cupy as cp\nimport numpy as np\nimport pandas as pd\nfrom cupyx.scipy.ndimage import gaussian_filter\nfrom cupyx.scipy.signal import filtfilt, iirnotch\nfrom cupyx.scipy.signal import spectrogram as cupyx_spectrogram\nfrom scipy.signal import filtfilt as scipy_filtfilt, butter as scipy_butter\n\ndef create_spectrogram_with_cupy(eeg_data, start, duration=50,\n                                 low_cut_freq=0.7, high_cut_freq=20, order_band=5,\n                                 spec_size_freq=267, spec_size_time=30,\n                                 nperseg=1500, noverlap=1483, nfft=2750,\n                                 sigma_gaussian=0.7,\n                                 mean_montage_names=4):\n    electrode_pair_name_locations = {'LL': ['Fp1', 'F7', 'T3', 'T5', 'O1'],\n                                     'RL': ['Fp2', 'F8', 'T4', 'T6', 'O2'],\n                                     'LP': ['Fp1', 'F3', 'C3', 'P3', 'O1'],\n                                     'RP': ['Fp2', 'F4', 'C4', 'P4', 'O2']}\n\n    # Filter specifications\n    nyquist_freq = 0.5 * 200\n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    # Bandpass and notch filter\n    # bandpass_coefficients = butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized], btype='band')\n    notch_coefficients = iirnotch(w0=60, Q=30, fs=200)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype='band')\n\n    spec_size = duration * 200\n    start = start * 200\n    real_start = start + (10_000 // 2) - (spec_size // 2)\n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    # Spectrogram parameters\n    fs = 200\n\n    if spec_size_freq <= 0 or spec_size_time <= 0:\n        freq_size = int((nfft // 2) / 5.15198) + 1\n        segments = int((spec_size - noverlap) / (nperseg - noverlap))\n    else:\n        freq_size = spec_size_freq\n        segments = spec_size_time\n\n    # Initialize spectrogram container\n    spectrogram = cp.zeros((freq_size, segments, 4), dtype='float32')\n\n    processed_eeg = {}\n\n    for i, (electrode_pair_name, electrode_locs) in enumerate(electrode_pair_name_locations.items()):\n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n        for j in range(4):\n            # Compute differential signals\n            signal = cp.array(eeg_data[electrode_locs[j]].values - eeg_data[electrode_locs[j + 1]].values)\n\n            # Handles NaNs \n            mean_signal = cp.nanmean(signal)\n            signal = cp.nan_to_num(signal, nan=mean_signal) if cp.isnan(signal).mean() < 1 else cp.zeros_like(signal)\n\n            # Filters bandpass and notch\n            signal_filtered = filtfilt(*notch_coefficients, signal)\n            signal_filtered = scipy_filtfilt(*sci_bandpass_coefficients, signal_filtered.get())  # HOTFIX\n\n            # GPU-accelerated spectrogram computation\n            frequencies, times, Sxx = cupyx_spectrogram(signal_filtered, fs, nperseg=nperseg, noverlap=noverlap,\n                                                        nfft=nfft)\n\n            # Filters frequency range \n            valid_freq = (frequencies >= 0.59) & (frequencies <= 20)\n            Sxx_filtered = Sxx[valid_freq, :]\n\n            # Logarithmic transformation and normalization using Cupy\n            spectrogram_slice = cp.clip(Sxx_filtered, cp.exp(-4), cp.exp(6))\n            spectrogram_slice = cp.log10(spectrogram_slice)\n\n            normalization_epsilon = 1e-6\n            mean = spectrogram_slice.mean(axis=(0, 1), keepdims=True)\n            std = spectrogram_slice.std(axis=(0, 1), keepdims=True)\n            spectrogram_slice = (spectrogram_slice - mean) / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[f'{electrode_locs[j]}_{electrode_locs[j + 1]}'] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        # AVERAGES THE 4 MONTAGE DIFFERENCES\n        if mean_montage_names > 0:\n            spectrogram[:, :, i] /= mean_montage_names\n\n    # Applies Gaussian filter and retrieves the spectrogram as a NumPy array using cupy.ndarray.get()\n    spec_numpy = gaussian_filter(spectrogram, sigma=sigma_gaussian).get() if sigma_gaussian > 0 else spectrogram.get()\n\n    # Filter EKG signal\n    ekg_signal_filtered = filtfilt(*notch_coefficients, cp.array(eeg_data[\"EKG\"].values))\n    processed_eeg['EKG'] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered.get())  # HOTFIX\n    return spec_numpy, processed_eeg\n```\n\n##Bonus\nAlso, after preparing 50 sec and 10 sec specs you can create the final image differently. So for example, rather than resizing it you can cut high frequencies like shown below and center left-right parts mirror-like.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Feefd98c6ccc8b8a1afb782d0952175fe%2F1.png?generation=1711557383664757&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ffeb52985983b67376f4aeef64742937e%2F2.png?generation=1711557395408198&alt=media)\n\nThough the originally proposed method used by Rafael brings the best score for us. You are welcome to share your ideas on how you combine the thing, but I am pretty much sure top teams did not leave Rafael's notebook unnoticed.\n\n##Outro\nHave fun and good luck, there is still a time to improve!",
      "votes": 52
    },
    {
      "id": 2719358,
      "postDate": "2024-03-27T16:46:00.527Z",
      "content": "<p>P.s. Let me know if you want me to post the rewritten Rafael's notebook for data preparation with both head cropping part and his original method.</p>",
      "rawMarkdown": "P.s. Let me know if you want me to post the rewritten Rafael's notebook for data preparation with both head cropping part and his original method.",
      "votes": 5,
      "replies": [
        {
          "id": 2719959,
          "postDate": "2024-03-28T03:18:50.807Z",
          "content": "<p>Thanks for sharing. I got a \"Runtime compilation failed error\" at the line mean_signal = cp.nanmean(signal). The environment was forked from <a href=\"https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\" target=\"_blank\">https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu</a>.</p>",
          "rawMarkdown": "Thanks for sharing. I got a \"Runtime compilation failed error\" at the line mean_signal = cp.nanmean(signal). The environment was forked from https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu.",
          "votes": 1,
          "replies": [
            {
              "id": 2719967,
              "postDate": "2024-03-28T03:25:51.157Z",
              "content": "<p>Did the original author notebook fail?</p>",
              "rawMarkdown": "Did the original author notebook fail?"
            },
            {
              "id": 2720184,
              "postDate": "2024-03-28T06:58:48.630Z",
              "content": "<p>There is no problem in completely copying the original author. I changed it a little bit, and the above error will appear when running on P100. However, in T4, the \"import cupy\" sentence will sometimes report an error, and sometimes it will not. I feel that the environment is not installed correctly. Can you share your installation steps, thanks.</p>",
              "rawMarkdown": "There is no problem in completely copying the original author. I changed it a little bit, and the above error will appear when running on P100. However, in T4, the \"import cupy\" sentence will sometimes report an error, and sometimes it will not. I feel that the environment is not installed correctly. Can you share your installation steps, thanks.",
              "votes": 1
            },
            {
              "id": 2721949,
              "postDate": "2024-03-29T10:14:22.033Z",
              "content": "<p>Maxwell or Pascal GPUs will be supported in CuPy v13.1.0. Therefore, an error (Runtime compilation failed error) will be raised on P100 at present. There is a similar <a href=\"https://github.com/cupy/cupy/issues/8260\" target=\"_blank\">issue</a> in CuPy's repo. <code>pip install git+https://github.com/cupy/cupy.git</code> can address this error.</p>",
              "rawMarkdown": "Maxwell or Pascal GPUs will be supported in CuPy v13.1.0. Therefore, an error (Runtime compilation failed error) will be raised on P100 at present. There is a similar [issue](https://github.com/cupy/cupy/issues/8260) in CuPy's repo. `pip install git+https://github.com/cupy/cupy.git` can address this error.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2728222,
          "postDate": "2024-04-02T06:24:51.837Z",
          "content": "<p>need help for head cropping part😭</p>",
          "rawMarkdown": "need help for head cropping part😭",
          "votes": 1,
          "replies": [
            {
              "id": 2740148,
              "postDate": "2024-04-07T16:11:40.410Z",
              "content": "<pre><code> ():\n    \n     cut_head:\n        single_channel_image1 = np.vstack([image_50s[..., i][:]  i  ()])\n        single_channel_image2 = np.vstack([image_10s[..., i][:]  i  ()])\n\n        \n        resized_image1 = cv2.resize(single_channel_image1, (, ), interpolation=cv2.INTER_AREA)\n        resized_image2 = cv2.resize(single_channel_image2, (, ), interpolation=cv2.INTER_AREA)\n        resized_image3 = cv2.resize(image_10m, (, ), interpolation=cv2.INTER_AREA)\n\n        \n        final_image = np.zeros((, ), dtype=np.float32)\n        final_image[:, :] = resized_image1\n        final_image[:, :] = resized_image2\n        final_image[:, :] = resized_image3  \n</code></pre>",
              "rawMarkdown": "```python\ndef create_final_image(image_50s, image_10s, image_10m, cut_head=True):\n    \"\"\"Combine three images into a single final image.\"\"\"\n    if cut_head:\n        single_channel_image1 = np.vstack([image_50s[..., i][:128] for i in range(4)])\n        single_channel_image2 = np.vstack([image_10s[..., i][:64] for i in range(4)])\n        \n        # Resize images to fit the final composition\n        resized_image1 = cv2.resize(single_channel_image1, (256, 512), interpolation=cv2.INTER_AREA)\n        resized_image2 = cv2.resize(single_channel_image2, (256, 256), interpolation=cv2.INTER_AREA)\n        resized_image3 = cv2.resize(image_10m, (256, 256), interpolation=cv2.INTER_AREA)\n        \n        # Create the final image and place the resized images accordingly\n        final_image = np.zeros((512, 512), dtype=np.float32)\n        final_image[0:512, 0:256] = resized_image1\n        final_image[0:256, 256:512] = resized_image2\n        final_image[256:512, 256:512] = resized_image3  \n```",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2735113,
      "postDate": "2024-04-04T14:31:11.253Z",
      "content": "<p>Hey friends, I made a bug fix in the model that resulted in an LB of 0.31. Below is the link to the discussion where I explain what happened :)</p>\n<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/491070\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/491070</a></p>",
      "rawMarkdown": "Hey friends, I made a bug fix in the model that resulted in an LB of 0.31. Below is the link to the discussion where I explain what happened :)\n\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/491070",
      "votes": 3,
      "replies": [
        {
          "id": 2735119,
          "postDate": "2024-04-04T14:40:42.450Z",
          "content": "<p>thank you Rafael</p>",
          "rawMarkdown": "thank you Rafael",
          "votes": 1
        }
      ]
    },
    {
      "id": 2747734,
      "postDate": "2024-04-12T03:40:18.713Z",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> cusignal is deprecated by NVIDIA. The RAPIDS team has contributed its content to cupy.</p>\n<p>This is to say that you are absolutely right! One should now use cupy to perform what cusignal was doing. NVIDIA developed cusignal to show that it is possible to perform signal processing on GPU. Moving it to cupy was the right way to make it both simpler to use, and more accessible.</p>",
      "rawMarkdown": "@sergiosaharovskiy cusignal is deprecated by NVIDIA. The RAPIDS team has contributed its content to cupy.\n\nThis is to say that you are absolutely right! One should now use cupy to perform what cusignal was doing. NVIDIA developed cusignal to show that it is possible to perform signal processing on GPU. Moving it to cupy was the right way to make it both simpler to use, and more accessible.",
      "votes": 4
    },
    {
      "id": 2734973,
      "postDate": "2024-04-04T12:56:39.023Z",
      "content": "<p>Thanks for the implementation. It does indeed work faster. But I was hoping that your and Rafael's approach would give a result close to 0.32lb. Hmm, maybe I'm doing something wrong)</p>",
      "rawMarkdown": "Thanks for the implementation. It does indeed work faster. But I was hoping that your and Rafael's approach would give a result close to 0.32lb. Hmm, maybe I'm doing something wrong)",
      "votes": 1,
      "replies": [
        {
          "id": 2734978,
          "postDate": "2024-04-04T13:02:59.463Z",
          "content": "<p>Well, I did not exactly copied the Rafael's pipeline. I have heavily borrowed his data preprocessing part, though for training I used the regular efficientnetb0 without any pseudo labeling and teacher-student model thing. And still two stage approach, hope it helps! </p>",
          "rawMarkdown": "Well, I did not exactly copied the Rafael's pipeline. I have heavily borrowed his data preprocessing part, though for training I used the regular efficientnetb0 without any pseudo labeling and teacher-student model thing. And still two stage approach, hope it helps! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2722528,
      "postDate": "2024-03-29T16:18:16.193Z",
      "content": "<p>Hi Sergey, may I ask which version of cupy are you installing? I tried multiple versions, either some of the following  are missing </p>\n<p>from cupyx.scipy.ndimage import gaussian_filter<br>\nfrom cupyx.scipy.signal import filtfilt, iirnotch<br>\nfrom cupyx.scipy.signal import spectrogram as cupyx_spectrogram</p>\n<p>or requiring some <code>lbnvrtc.so</code> stuff, very annoying…</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hi Sergey, may I ask which version of cupy are you installing? I tried multiple versions, either some of the following  are missing \n\nfrom cupyx.scipy.ndimage import gaussian_filter\nfrom cupyx.scipy.signal import filtfilt, iirnotch\nfrom cupyx.scipy.signal import spectrogram as cupyx_spectrogram\n\nor requiring some `lbnvrtc.so` stuff, very annoying...\n\nThank you!",
      "votes": 1,
      "replies": [
        {
          "id": 2722590,
          "postDate": "2024-03-29T17:03:19.683Z",
          "content": "<p>Hi Qui, </p>\n<p>Locally I use the following:</p>\n<p>cupy-cuda11x==13.0.0<br>\nnvidia-cuda-nvrtc-cu11==11.7.99<br>\nnvidia-cuda-cupti-cu11==11.7.101<br>\nnvidia-cuda-runtime-cu11==11.7.99</p>",
          "rawMarkdown": "Hi Qui, \n\nLocally I use the following:\n\ncupy-cuda11x==13.0.0\nnvidia-cuda-nvrtc-cu11==11.7.99\nnvidia-cuda-cupti-cu11==11.7.101\nnvidia-cuda-runtime-cu11==11.7.99\n",
          "votes": 1,
          "replies": [
            {
              "id": 2723610,
              "postDate": "2024-03-30T10:40:20.933Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2725289,
              "postDate": "2024-03-31T13:10:27.787Z",
              "content": "<p>Hi Sergey, great thanks for your information! <br>\nI finally can import all the packages used in your code! <br>\nPreviously, I tried multiple version of cupy and never made it. <br>\nThank you!</p>",
              "rawMarkdown": "Hi Sergey, great thanks for your information! \nI finally can import all the packages used in your code! \nPreviously, I tried multiple version of cupy and never made it. \nThank you!",
              "votes": 1
            },
            {
              "id": 2725468,
              "postDate": "2024-03-31T15:05:19.683Z",
              "content": "<p>Unfortunately, i can import pacakges sucessfully but cannot actually run the function … have no idea about what's happening…</p>\n<p><code>RuntimeError: CuPy failed to load libnvrtc.so.11.2: OSError: libnvrtc.so.11.2: cannot open shared object file: No such file or directory</code></p>\n<p>My system cuda is <code>cuda 11.7</code> and my pytorch version is <code>torch==2.0.1+cu118</code>, maybe some conflicts on this? </p>\n<p>Also, does your code run in Kaggle notebook? I tried to run the function on Kaggle and got a different error:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552698%2F4b7b1ed2dca8392cfbf6140d1cbdaa4d%2FCupyBug2.png?generation=1711897836487915&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552698%2Fabf57a8c12eae4ad4a6627e1ba2fd11b%2FCupyBug.jpg?generation=1711897771631898&amp;alt=media\"></p>\n<p>Thank you!</p>",
              "rawMarkdown": "Unfortunately, i can import pacakges sucessfully but cannot actually run the function ... have no idea about what's happening...\n\n`RuntimeError: CuPy failed to load libnvrtc.so.11.2: OSError: libnvrtc.so.11.2: cannot open shared object file: No such file or directory`\n\nMy system cuda is `cuda 11.7` and my pytorch version is `torch==2.0.1+cu118`, maybe some conflicts on this? \n\nAlso, does your code run in Kaggle notebook? I tried to run the function on Kaggle and got a different error:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552698%2F4b7b1ed2dca8392cfbf6140d1cbdaa4d%2FCupyBug2.png?generation=1711897836487915&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552698%2Fabf57a8c12eae4ad4a6627e1ba2fd11b%2FCupyBug.jpg?generation=1711897771631898&alt=media)\n\nThank you!"
            },
            {
              "id": 2725531,
              "postDate": "2024-03-31T16:00:07.457Z",
              "content": "<p>switch to T4 gpu?</p>",
              "rawMarkdown": "switch to T4 gpu?"
            },
            {
              "id": 2740811,
              "postDate": "2024-04-08T00:39:14.490Z",
              "content": "<p>Hi Sergey, sorry for the late reply. I finally got it work on Kaggle T4 GPUs. However, I tried to print out both CUDA versions of T4 &amp; P100, they were both CUDA12.1 (or something). But it failed to compile on P100, which I don't know why. Maybe related to hardware architecture compatibility?</p>",
              "rawMarkdown": "Hi Sergey, sorry for the late reply. I finally got it work on Kaggle T4 GPUs. However, I tried to print out both CUDA versions of T4 & P100, they were both CUDA12.1 (or something). But it failed to compile on P100, which I don't know why. Maybe related to hardware architecture compatibility?\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2721690,
      "postDate": "2024-03-29T06:41:25.967Z",
      "content": "<p>Thanks for sharing!<br>\nDid you do use the same preprocessing’s parameter with the Rafael’s notebook?</p>",
      "rawMarkdown": "Thanks for sharing!\nDid you do use the same preprocessing’s parameter with the Rafael’s notebook?",
      "votes": 1,
      "replies": [
        {
          "id": 2721995,
          "postDate": "2024-03-29T10:41:21.160Z",
          "content": "<p>yes, I use the same parameters</p>",
          "rawMarkdown": "yes, I use the same parameters",
          "votes": 1
        }
      ]
    },
    {
      "id": 2720985,
      "postDate": "2024-03-28T17:28:33.613Z",
      "content": "<p>Thanks for sharing! Did you do any preprocessing to <code>eeg_data</code> before passing it into the function? With <code>eeg_data</code> being <code>pd.read_parquet(...)</code>, and calling your function, I get the error: <code>ValueError: operands could not be broadcast together with shapes (267, 30) (267, 501) (267, 30)</code> on line <code>spectrogram[i, :, :] += spectrogram_slice</code>.</p>\n<p>Not sure if you've trimmed the loaded <code>eeg_data</code> before calling that function?</p>",
      "rawMarkdown": "Thanks for sharing! Did you do any preprocessing to `eeg_data` before passing it into the function? With `eeg_data` being `pd.read_parquet(...)`, and calling your function, I get the error: `ValueError: operands could not be broadcast together with shapes (267, 30) (267, 501) (267, 30)` on line `spectrogram[i, :, :] += spectrogram_slice`.\n\nNot sure if you've trimmed the loaded `eeg_data` before calling that function?",
      "votes": 1,
      "replies": [
        {
          "id": 2721047,
          "postDate": "2024-03-28T18:05:23.353Z",
          "content": "<p>update: can be fixed by calculating the shape on the fly with <code>segments = int((spec_size - noverlap) / (nperseg - noverlap))</code> instead of specifying as an argument</p>",
          "rawMarkdown": "update: can be fixed by calculating the shape on the fly with `segments = int((spec_size - noverlap) / (nperseg - noverlap))` instead of specifying as an argument",
          "votes": 2
        }
      ]
    },
    {
      "id": 2719621,
      "postDate": "2024-03-27T20:36:55.973Z",
      "content": "<p>Thank you for the mention, Sergey. I thought your code optimization was quite good. Indeed, installing Rapids can sometimes be slow and end up being a hassle, but this way, it's much easier to use.</p>",
      "rawMarkdown": "Thank you for the mention, Sergey. I thought your code optimization was quite good. Indeed, installing Rapids can sometimes be slow and end up being a hassle, but this way, it's much easier to use.",
      "votes": 1,
      "replies": [
        {
          "id": 2719744,
          "postDate": "2024-03-27T22:24:22.020Z",
          "content": "<p>Yes, thank you Rafael. I always try to give a credit where the credit is due. I believe we gonna see this data preparation pipeline in the top solutions.</p>",
          "rawMarkdown": "Yes, thank you Rafael. I always try to give a credit where the credit is due. I believe we gonna see this data preparation pipeline in the top solutions.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2736821,
      "postDate": "2024-04-05T12:19:21.250Z",
      "content": "<p>Throw multiple error with me outside KAGGLE, the fix from my point of view is below, however it may be experience low performance or different results.</p>\n<pre><code>def create_spectrogram_with_cupy(eeg_data, start, duration=,\n                                 low_cut_freq=, high_cut_freq=, order_band=,\n                                 spec_size_freq=, spec_size_time=,\n                                 nperseg=, noverlap=, nfft=,\n                                 sigma_gaussian=,\n                                 mean_montage_names=):\n    electrode_pair_name_locations = {: [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ]}\n\n    # Filter specifications\n    nyquist_freq =  * \n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    # Bandpass  notch filter\n    # bandpass_coefficients = butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized], btype=)\n    notch_coefficients = iirnotch(w0=, Q=, fs=)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype=)\n\n    spec_size = duration * \n    start = start * \n    real_start = start + (_000 \n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    # Spectrogram \n    fs \n\n    if \n        freq_size  //  / ) + \n        segments = int((spec_size - noverlap) / \n    else:\n        freq_size \n        segments \n\n    # Initialize \n    spectrogram \n\n    processed_eeg \n\n    for \n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n        for \n            # Compute \n            signal \n\n            # Handles  \n            mean_signal \n            signal \n\n            signal_np \n\n            # Filters \n            signal_filtered \n            signal_filtered \n\n            # GPU-accelerated \n            frequencies, times, Sxx \n                                                        nfft=nfft)\n\n            # Filters  \n            valid_freq \n            Sxx_filtered \n\n            # Logarithmic \n            spectrogram_slice \n            spectrogram_slice \n\n            normalization_epsilon \n            mean \n            std \n            spectrogram_slice  / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[f] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        # AVERAGES THE  MONTAGE DIFFERENCES\n         mean_montage_names &gt; :\n            spectrogram[:, :, i] /\n\n    # Applies \n    spec_numpy \n\n    # Filter \n    ekg_np \n    ekg_signal_filtered \n    processed_eeg[] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered)  # HOTFIX\n    return \n</code></pre>",
      "rawMarkdown": "Throw multiple error with me outside KAGGLE, the fix from my point of view is below, however it may be experience low performance or different results.\n\n```\ndef create_spectrogram_with_cupy(eeg_data, start, duration=50,\n                                 low_cut_freq=0.7, high_cut_freq=20, order_band=5,\n                                 spec_size_freq=267, spec_size_time=30,\n                                 nperseg=1500, noverlap=1483, nfft=2750,\n                                 sigma_gaussian=0.7,\n                                 mean_montage_names=4):\n    electrode_pair_name_locations = {'LL': ['Fp1', 'F7', 'T3', 'T5', 'O1'],\n                                     'RL': ['Fp2', 'F8', 'T4', 'T6', 'O2'],\n                                     'LP': ['Fp1', 'F3', 'C3', 'P3', 'O1'],\n                                     'RP': ['Fp2', 'F4', 'C4', 'P4', 'O2']}\n\n    # Filter specifications\n    nyquist_freq = 0.5 * 200\n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    # Bandpass and notch filter\n    # bandpass_coefficients = butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized], btype='band')\n    notch_coefficients = iirnotch(w0=60, Q=30, fs=200)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype='band')\n\n    spec_size = duration * 200\n    start = start * 200\n    real_start = start + (10_000 // 2) - (spec_size // 2)\n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    # Spectrogram parameters\n    fs = 200\n\n    if spec_size_freq <= 0 or spec_size_time <= 0:\n        freq_size = int((nfft // 2) / 5.15198) + 1\n        segments = int((spec_size - noverlap) / (nperseg - noverlap))\n    else:\n        freq_size = spec_size_freq\n        segments = spec_size_time\n\n    # Initialize spectrogram container\n    spectrogram = cp.zeros((freq_size, segments, 4), dtype='float32')\n\n    processed_eeg = {}\n\n    for i, (electrode_pair_name, electrode_locs) in enumerate(electrode_pair_name_locations.items()):\n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n        for j in range(4):\n            # Compute differential signals\n            signal = cp.array(eeg_data[electrode_locs[j]].values - eeg_data[electrode_locs[j + 1]].values)\n\n            # Handles NaNs \n            mean_signal = cp.nanmean(signal)\n            signal = cp.nan_to_num(signal, nan=mean_signal) if cp.isnan(signal).mean() < 1 else cp.zeros_like(signal)\n\n            signal_np = signal.get()\n\n            # Filters bandpass and notch\n            signal_filtered = filtfilt(*notch_coefficients, signal_np)\n            signal_filtered = scipy_filtfilt(*sci_bandpass_coefficients, signal_filtered)  # HOTFIX\n\n            # GPU-accelerated spectrogram computation\n            frequencies, times, Sxx = cupyx_spectrogram(signal_filtered, fs, nperseg=nperseg, noverlap=noverlap,\n                                                        nfft=nfft)\n\n            # Filters frequency range \n            valid_freq = (frequencies >= 0.59) & (frequencies <= 20)\n            Sxx_filtered = Sxx[valid_freq, :]\n\n            # Logarithmic transformation and normalization using Cupy\n            spectrogram_slice = cp.clip(Sxx_filtered, cp.exp(-4), cp.exp(6))\n            spectrogram_slice = cp.log10(spectrogram_slice)\n\n            normalization_epsilon = 1e-6\n            mean = spectrogram_slice.mean(axis=(0, 1), keepdims=True)\n            std = spectrogram_slice.std(axis=(0, 1), keepdims=True)\n            spectrogram_slice = (spectrogram_slice - mean) / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[f'{electrode_locs[j]}_{electrode_locs[j + 1]}'] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        # AVERAGES THE 4 MONTAGE DIFFERENCES\n        if mean_montage_names > 0:\n            spectrogram[:, :, i] /= mean_montage_names\n\n    # Applies Gaussian filter and retrieves the spectrogram as a NumPy array using cupy.ndarray.get()\n    spec_numpy = gaussian_filter(spectrogram, sigma=sigma_gaussian).get() if sigma_gaussian > 0 else spectrogram.get()\n\n    # Filter EKG signal\n    ekg_np = cp.array(eeg_data[\"EKG\"].values).get()\n    ekg_signal_filtered = filtfilt(*notch_coefficients, ekg_np )\n    processed_eeg['EKG'] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered)  # HOTFIX\n    return spec_numpy, processed_eeg\n\n```",
      "votes": 2
    },
    {
      "id": 2720308,
      "postDate": "2024-03-28T08:42:57.593Z",
      "content": "<p>Thanks for sharing.<br>\nMy experiment is also going well with Rafael's data.<br>\nDo you use masking augmentation ? </p>",
      "rawMarkdown": "Thanks for sharing.\nMy experiment is also going well with Rafael's data.\nDo you use masking augmentation ? \n ",
      "votes": 2,
      "replies": [
        {
          "id": 2720392,
          "postDate": "2024-03-28T10:06:24.250Z",
          "content": "<p>I dont, hflip only works for me.</p>",
          "rawMarkdown": "I dont, hflip only works for me.",
          "votes": 3,
          "replies": [
            {
              "id": 2720846,
              "postDate": "2024-03-28T16:15:13.543Z",
              "content": "<blockquote>\n  <p>Honestly, it is how we get our current 0.28 on the leaderboard.</p>\n</blockquote>\n<p>Do you mean that you got 0.28 with single model trained by  Rafael's data ? </p>",
              "rawMarkdown": ">Honestly, it is how we get our current 0.28 on the leaderboard.\n\nDo you mean that you got 0.28 with single model trained by  Rafael's data ? "
            },
            {
              "id": 2720855,
              "postDate": "2024-03-28T16:24:41.560Z",
              "content": "<p>single model is around 0.3</p>",
              "rawMarkdown": "single model is around 0.3"
            },
            {
              "id": 2720867,
              "postDate": "2024-03-28T16:32:08.783Z",
              "content": "<p>my notebook with pytorch model ran successfully when run&amp;save, but failed when submission, <strong>Notebook Threw Exception</strong><br>\nNothing change about data processing but loading and inferring..<br>\nIs there any advice with this problem? </p>",
              "rawMarkdown": "my notebook with pytorch model ran successfully when run&save, but failed when submission, **Notebook Threw Exception**\nNothing change about data processing but loading and inferring..\nIs there any advice with this problem? "
            },
            {
              "id": 2721452,
              "postDate": "2024-03-29T03:04:49.193Z",
              "content": "<p>I solved the problem by using numpy and scipy instead of cupy and cusignal.<br>\nI rewrite data generation script in my local with numpy, scipy and multiprocessing.<br>\nIt takes 30 minutes to generate data in my local and  about 1 hour to submit successfully.</p>",
              "rawMarkdown": "I solved the problem by using numpy and scipy instead of cupy and cusignal.\nI rewrite data generation script in my local with numpy, scipy and multiprocessing.\nIt takes 30 minutes to generate data in my local and  about 1 hour to submit successfully."
            },
            {
              "id": 2721994,
              "postDate": "2024-03-29T10:40:48.097Z",
              "content": "<p>my inference runs on T4 with cupy, not problem there</p>",
              "rawMarkdown": "my inference runs on T4 with cupy, not problem there",
              "votes": 1
            },
            {
              "id": 2722315,
              "postDate": "2024-03-29T14:09:31.360Z",
              "content": "<p>Can you tell me how to use cupy inference on T4? My inference using cupy is very slow on T4.</p>",
              "rawMarkdown": "Can you tell me how to use cupy inference on T4? My inference using cupy is very slow on T4."
            },
            {
              "id": 2730342,
              "postDate": "2024-04-02T16:54:24.910Z",
              "content": "<p>I got the same score. LB0.30</p>",
              "rawMarkdown": "I got the same score. LB0.30",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2723327,
      "postDate": "2024-03-30T05:50:08.103Z",
      "content": "<p>Hi Qui,</p>\n<p>Locally I use the following:</p>\n<p>cupy-cuda11x==13.0.0<br>\nnvidia-cuda-nvrtc-cu11==11.7.99<br>\nnvidia-cuda-cupti-cu11==11.7.101<br>\nnvidia-cuda-runtime-cu11==11.7.99</p>",
      "rawMarkdown": "Hi Qui,\n\nLocally I use the following:\n\ncupy-cuda11x==13.0.0\nnvidia-cuda-nvrtc-cu11==11.7.99\nnvidia-cuda-cupti-cu11==11.7.101\nnvidia-cuda-runtime-cu11==11.7.99",
      "votes": -1
    },
    {
      "id": 2720795,
      "postDate": "2024-03-28T15:37:34.880Z",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> with Melspectrogram I know I can restrict output spectrogram to a certain size, but cupy signal spectrogram seems to be a little tricky. Is there any way to control the output size of the spectrogram?</p>",
      "rawMarkdown": "@sergiosaharovskiy with Melspectrogram I know I can restrict output spectrogram to a certain size, but cupy signal spectrogram seems to be a little tricky. Is there any way to control the output size of the spectrogram?"
    },
    {
      "id": 2719895,
      "postDate": "2024-03-28T02:04:54.273Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2719900,
          "postDate": "2024-03-28T02:18:16.633Z",
          "content": "<p>the improvement is to replace everything from cusignal + scipy to cupy only.</p>",
          "rawMarkdown": "the improvement is to replace everything from cusignal + scipy to cupy only.",
          "replies": [
            {
              "id": 2720181,
              "postDate": "2024-03-28T06:57:15.363Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2720602,
              "postDate": "2024-03-28T13:24:17.190Z",
              "content": "<p>I used 2 gpus, it took me~15 minute to prepare the whole thing.</p>",
              "rawMarkdown": "I used 2 gpus, it took me~15 minute to prepare the whole thing."
            },
            {
              "id": 2720686,
              "postDate": "2024-03-28T14:18:08.727Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10391122%2F36d06dcd61a96a1186321918358a1c69%2F283.png?generation=1711635510324705&amp;alt=media\"><br>\nI modified my code to follow yours above, but it took longer than the original using CuSignal. In addition, GPU utilization is very low. Do you have any suggestions to tackle this issue?</p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10391122%2F36d06dcd61a96a1186321918358a1c69%2F283.png?generation=1711635510324705&alt=media)\nI modified my code to follow yours above, but it took longer than the original using CuSignal. In addition, GPU utilization is very low. Do you have any suggestions to tackle this issue?",
              "votes": 1
            },
            {
              "id": 2720733,
              "postDate": "2024-03-28T14:33:58.990Z",
              "content": "<p>try to parallel it or try to prepare half of the data on one device and another half on another</p>\n<pre><code> joblib  Parallel, delayed\nnum_cores = \nParallel(n_jobs=num_cores)(delayed(process_eeg)(i)  i  tqdm(((train_df)), desc=))\n</code></pre>",
              "rawMarkdown": "  try to parallel it or try to prepare half of the data on one device and another half on another\n```python\n from joblib import Parallel, delayed\n num_cores = 4\n Parallel(n_jobs=num_cores)(delayed(process_eeg)(i) for i in tqdm(range(len(train_df)), desc=\"Processing EEGs\"))\n```",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2723867,
      "postDate": "2024-03-30T14:33:30.190Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 2719358,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2024-03-27T16:46:00.527000",
      "content": "<p>P.s. Let me know if you want me to post the rewritten Rafael's notebook for data preparation with both head cropping part and his original method.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2719959,
          "author_name": "ynhuhu",
          "author_url": "",
          "post_date": "2024-03-28T03:18:50.807000",
          "content": "<p>Thanks for sharing. I got a \"Runtime compilation failed error\" at the line mean_signal = cp.nanmean(signal). The environment was forked from <a href=\"https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu\" target=\"_blank\">https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu</a>.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2719967,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-03-28T03:25:51.157000",
              "content": "<p>Did the original author notebook fail?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2720184,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2024-03-28T06:58:48.630000",
              "content": "<p>There is no problem in completely copying the original author. I changed it a little bit, and the above error will appear when running on P100. However, in T4, the \"import cupy\" sentence will sometimes report an error, and sometimes it will not. I feel that the environment is not installed correctly. Can you share your installation steps, thanks.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2721949,
              "author_name": "Koolo",
              "author_url": "",
              "post_date": "2024-03-29T10:14:22.033000",
              "content": "<p>Maxwell or Pascal GPUs will be supported in CuPy v13.1.0. Therefore, an error (Runtime compilation failed error) will be raised on P100 at present. There is a similar <a href=\"https://github.com/cupy/cupy/issues/8260\" target=\"_blank\">issue</a> in CuPy's repo. <code>pip install git+https://github.com/cupy/cupy.git</code> can address this error.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2728222,
          "author_name": "pky",
          "author_url": "",
          "post_date": "2024-04-02T06:24:51.837000",
          "content": "<p>need help for head cropping part😭</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2740148,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-04-07T16:11:40.410000",
              "content": "<pre><code> ():\n    \n     cut_head:\n        single_channel_image1 = np.vstack([image_50s[..., i][:]  i  ()])\n        single_channel_image2 = np.vstack([image_10s[..., i][:]  i  ()])\n\n        \n        resized_image1 = cv2.resize(single_channel_image1, (, ), interpolation=cv2.INTER_AREA)\n        resized_image2 = cv2.resize(single_channel_image2, (, ), interpolation=cv2.INTER_AREA)\n        resized_image3 = cv2.resize(image_10m, (, ), interpolation=cv2.INTER_AREA)\n\n        \n        final_image = np.zeros((, ), dtype=np.float32)\n        final_image[:, :] = resized_image1\n        final_image[:, :] = resized_image2\n        final_image[:, :] = resized_image3  \n</code></pre>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2735113,
      "author_name": "Rafael Zimmermann",
      "author_url": "",
      "post_date": "2024-04-04T14:31:11.253000",
      "content": "<p>Hey friends, I made a bug fix in the model that resulted in an LB of 0.31. Below is the link to the discussion where I explain what happened :)</p>\n<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/491070\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/491070</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2735119,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-04-04T14:40:42.450000",
          "content": "<p>thank you Rafael</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2747734,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-04-12T03:40:18.713000",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> cusignal is deprecated by NVIDIA. The RAPIDS team has contributed its content to cupy.</p>\n<p>This is to say that you are absolutely right! One should now use cupy to perform what cusignal was doing. NVIDIA developed cusignal to show that it is possible to perform signal processing on GPU. Moving it to cupy was the right way to make it both simpler to use, and more accessible.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2734973,
      "author_name": "prince_lvov",
      "author_url": "",
      "post_date": "2024-04-04T12:56:39.023000",
      "content": "<p>Thanks for the implementation. It does indeed work faster. But I was hoping that your and Rafael's approach would give a result close to 0.32lb. Hmm, maybe I'm doing something wrong)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2734978,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-04-04T13:02:59.463000",
          "content": "<p>Well, I did not exactly copied the Rafael's pipeline. I have heavily borrowed his data preprocessing part, though for training I used the regular efficientnetb0 without any pseudo labeling and teacher-student model thing. And still two stage approach, hope it helps! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2722528,
      "author_name": "豆柴金鯱",
      "author_url": "",
      "post_date": "2024-03-29T16:18:16.193000",
      "content": "<p>Hi Sergey, may I ask which version of cupy are you installing? I tried multiple versions, either some of the following  are missing </p>\n<p>from cupyx.scipy.ndimage import gaussian_filter<br>\nfrom cupyx.scipy.signal import filtfilt, iirnotch<br>\nfrom cupyx.scipy.signal import spectrogram as cupyx_spectrogram</p>\n<p>or requiring some <code>lbnvrtc.so</code> stuff, very annoying…</p>\n<p>Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2722590,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-29T17:03:19.683000",
          "content": "<p>Hi Qui, </p>\n<p>Locally I use the following:</p>\n<p>cupy-cuda11x==13.0.0<br>\nnvidia-cuda-nvrtc-cu11==11.7.99<br>\nnvidia-cuda-cupti-cu11==11.7.101<br>\nnvidia-cuda-runtime-cu11==11.7.99</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2723610,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-30T10:40:20.933000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2725289,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2024-03-31T13:10:27.787000",
              "content": "<p>Hi Sergey, great thanks for your information! <br>\nI finally can import all the packages used in your code! <br>\nPreviously, I tried multiple version of cupy and never made it. <br>\nThank you!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2725468,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2024-03-31T15:05:19.683000",
              "content": "<p>Unfortunately, i can import pacakges sucessfully but cannot actually run the function … have no idea about what's happening…</p>\n<p><code>RuntimeError: CuPy failed to load libnvrtc.so.11.2: OSError: libnvrtc.so.11.2: cannot open shared object file: No such file or directory</code></p>\n<p>My system cuda is <code>cuda 11.7</code> and my pytorch version is <code>torch==2.0.1+cu118</code>, maybe some conflicts on this? </p>\n<p>Also, does your code run in Kaggle notebook? I tried to run the function on Kaggle and got a different error:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552698%2F4b7b1ed2dca8392cfbf6140d1cbdaa4d%2FCupyBug2.png?generation=1711897836487915&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552698%2Fabf57a8c12eae4ad4a6627e1ba2fd11b%2FCupyBug.jpg?generation=1711897771631898&amp;alt=media\"></p>\n<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2725531,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-03-31T16:00:07.457000",
              "content": "<p>switch to T4 gpu?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2740811,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2024-04-08T00:39:14.490000",
              "content": "<p>Hi Sergey, sorry for the late reply. I finally got it work on Kaggle T4 GPUs. However, I tried to print out both CUDA versions of T4 &amp; P100, they were both CUDA12.1 (or something). But it failed to compile on P100, which I don't know why. Maybe related to hardware architecture compatibility?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2721690,
      "author_name": "MBOOK",
      "author_url": "",
      "post_date": "2024-03-29T06:41:25.967000",
      "content": "<p>Thanks for sharing!<br>\nDid you do use the same preprocessing’s parameter with the Rafael’s notebook?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2721995,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-29T10:41:21.160000",
          "content": "<p>yes, I use the same parameters</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2720985,
      "author_name": "Mandeep",
      "author_url": "",
      "post_date": "2024-03-28T17:28:33.613000",
      "content": "<p>Thanks for sharing! Did you do any preprocessing to <code>eeg_data</code> before passing it into the function? With <code>eeg_data</code> being <code>pd.read_parquet(...)</code>, and calling your function, I get the error: <code>ValueError: operands could not be broadcast together with shapes (267, 30) (267, 501) (267, 30)</code> on line <code>spectrogram[i, :, :] += spectrogram_slice</code>.</p>\n<p>Not sure if you've trimmed the loaded <code>eeg_data</code> before calling that function?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2721047,
          "author_name": "Mandeep",
          "author_url": "",
          "post_date": "2024-03-28T18:05:23.353000",
          "content": "<p>update: can be fixed by calculating the shape on the fly with <code>segments = int((spec_size - noverlap) / (nperseg - noverlap))</code> instead of specifying as an argument</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2719621,
      "author_name": "Rafael Zimmermann",
      "author_url": "",
      "post_date": "2024-03-27T20:36:55.973000",
      "content": "<p>Thank you for the mention, Sergey. I thought your code optimization was quite good. Indeed, installing Rapids can sometimes be slow and end up being a hassle, but this way, it's much easier to use.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2719744,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-27T22:24:22.020000",
          "content": "<p>Yes, thank you Rafael. I always try to give a credit where the credit is due. I believe we gonna see this data preparation pipeline in the top solutions.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2736821,
      "author_name": "Mahmoud Elshahed",
      "author_url": "",
      "post_date": "2024-04-05T12:19:21.250000",
      "content": "<p>Throw multiple error with me outside KAGGLE, the fix from my point of view is below, however it may be experience low performance or different results.</p>\n<pre><code>def create_spectrogram_with_cupy(eeg_data, start, duration=,\n                                 low_cut_freq=, high_cut_freq=, order_band=,\n                                 spec_size_freq=, spec_size_time=,\n                                 nperseg=, noverlap=, nfft=,\n                                 sigma_gaussian=,\n                                 mean_montage_names=):\n    electrode_pair_name_locations = {: [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ],\n                                     : [, , , , ]}\n\n    # Filter specifications\n    nyquist_freq =  * \n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    # Bandpass  notch filter\n    # bandpass_coefficients = butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized], btype=)\n    notch_coefficients = iirnotch(w0=, Q=, fs=)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype=)\n\n    spec_size = duration * \n    start = start * \n    real_start = start + (_000 \n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    # Spectrogram \n    fs \n\n    if \n        freq_size  //  / ) + \n        segments = int((spec_size - noverlap) / \n    else:\n        freq_size \n        segments \n\n    # Initialize \n    spectrogram \n\n    processed_eeg \n\n    for \n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n        for \n            # Compute \n            signal \n\n            # Handles  \n            mean_signal \n            signal \n\n            signal_np \n\n            # Filters \n            signal_filtered \n            signal_filtered \n\n            # GPU-accelerated \n            frequencies, times, Sxx \n                                                        nfft=nfft)\n\n            # Filters  \n            valid_freq \n            Sxx_filtered \n\n            # Logarithmic \n            spectrogram_slice \n            spectrogram_slice \n\n            normalization_epsilon \n            mean \n            std \n            spectrogram_slice  / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[f] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        # AVERAGES THE  MONTAGE DIFFERENCES\n         mean_montage_names &gt; :\n            spectrogram[:, :, i] /\n\n    # Applies \n    spec_numpy \n\n    # Filter \n    ekg_np \n    ekg_signal_filtered \n    processed_eeg[] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered)  # HOTFIX\n    return \n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2720308,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-03-28T08:42:57.593000",
      "content": "<p>Thanks for sharing.<br>\nMy experiment is also going well with Rafael's data.<br>\nDo you use masking augmentation ? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2720392,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-28T10:06:24.250000",
          "content": "<p>I dont, hflip only works for me.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2720846,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-03-28T16:15:13.543000",
              "content": "<blockquote>\n  <p>Honestly, it is how we get our current 0.28 on the leaderboard.</p>\n</blockquote>\n<p>Do you mean that you got 0.28 with single model trained by  Rafael's data ? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2720855,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-03-28T16:24:41.560000",
              "content": "<p>single model is around 0.3</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2720867,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-03-28T16:32:08.783000",
              "content": "<p>my notebook with pytorch model ran successfully when run&amp;save, but failed when submission, <strong>Notebook Threw Exception</strong><br>\nNothing change about data processing but loading and inferring..<br>\nIs there any advice with this problem? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2721452,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-03-29T03:04:49.193000",
              "content": "<p>I solved the problem by using numpy and scipy instead of cupy and cusignal.<br>\nI rewrite data generation script in my local with numpy, scipy and multiprocessing.<br>\nIt takes 30 minutes to generate data in my local and  about 1 hour to submit successfully.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2721994,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-03-29T10:40:48.097000",
              "content": "<p>my inference runs on T4 with cupy, not problem there</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2722315,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2024-03-29T14:09:31.360000",
              "content": "<p>Can you tell me how to use cupy inference on T4? My inference using cupy is very slow on T4.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2730342,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-02T16:54:24.910000",
              "content": "<p>I got the same score. LB0.30</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2723327,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-03-30T05:50:08.103000",
      "content": "<p>Hi Qui,</p>\n<p>Locally I use the following:</p>\n<p>cupy-cuda11x==13.0.0<br>\nnvidia-cuda-nvrtc-cu11==11.7.99<br>\nnvidia-cuda-cupti-cu11==11.7.101<br>\nnvidia-cuda-runtime-cu11==11.7.99</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2720795,
      "author_name": "Roy Wei",
      "author_url": "",
      "post_date": "2024-03-28T15:37:34.880000",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> with Melspectrogram I know I can restrict output spectrogram to a certain size, but cupy signal spectrogram seems to be a little tricky. Is there any way to control the output size of the spectrogram?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2719895,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-28T02:04:54.273000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2719900,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-28T02:18:16.633000",
          "content": "<p>the improvement is to replace everything from cusignal + scipy to cupy only.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2720181,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-28T06:57:15.363000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2720602,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-03-28T13:24:17.190000",
              "content": "<p>I used 2 gpus, it took me~15 minute to prepare the whole thing.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2720686,
              "author_name": "Dive Deeper",
              "author_url": "",
              "post_date": "2024-03-28T14:18:08.727000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10391122%2F36d06dcd61a96a1186321918358a1c69%2F283.png?generation=1711635510324705&amp;alt=media\"><br>\nI modified my code to follow yours above, but it took longer than the original using CuSignal. In addition, GPU utilization is very low. Do you have any suggestions to tackle this issue?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2720733,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-03-28T14:33:58.990000",
              "content": "<p>try to parallel it or try to prepare half of the data on one device and another half on another</p>\n<pre><code> joblib  Parallel, delayed\nnum_cores = \nParallel(n_jobs=num_cores)(delayed(process_eeg)(i)  i  tqdm(((train_df)), desc=))\n</code></pre>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2723867,
      "author_name": "Jiadong Zhu",
      "author_url": "",
      "post_date": "2024-03-30T14:33:30.190000",
      "content": "<p>thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2719349": "Hi,\n\n##Intro\nFirst things first, I'd like to say thank you to @rafaelzimmermann1 and his team for data preparation idea posted [in this notebook](https://www.kaggle.com/code/rafaelzimmermann1/hms-spectrogram-creation-using-gpu). Honestly, it is how we get our current **0.28** on the leaderboard. Though, I'd like to offer some improvements to the main script function `create_spectrogram_with_cusignal`. \n\n##The script\nRafael creates spectrogram by using Rapids, namely cuSignal, which might be a bit painful to install and keepup with versions. The exact same code can be rewritten  by using preinstalled cupy library (hasstle free solution). Here is the whole thing:\n\n```python\n# HOTFIX for cupy.signal.filtfilt with bandpass_coefficients cannot find the reason...\n# It produces different results than scipy.signal.filtfilt with the same filter coefficients.\n# signal_filtered = filtfilt(*bandpass_coefficients, signal_filtered)\n\n\nimport cupy as cp\nimport numpy as np\nimport pandas as pd\nfrom cupyx.scipy.ndimage import gaussian_filter\nfrom cupyx.scipy.signal import filtfilt, iirnotch\nfrom cupyx.scipy.signal import spectrogram as cupyx_spectrogram\nfrom scipy.signal import filtfilt as scipy_filtfilt, butter as scipy_butter\n\ndef create_spectrogram_with_cupy(eeg_data, start, duration=50,\n                                 low_cut_freq=0.7, high_cut_freq=20, order_band=5,\n                                 spec_size_freq=267, spec_size_time=30,\n                                 nperseg=1500, noverlap=1483, nfft=2750,\n                                 sigma_gaussian=0.7,\n                                 mean_montage_names=4):\n    electrode_pair_name_locations = {'LL': ['Fp1', 'F7', 'T3', 'T5', 'O1'],\n                                     'RL': ['Fp2', 'F8', 'T4', 'T6', 'O2'],\n                                     'LP': ['Fp1', 'F3', 'C3', 'P3', 'O1'],\n                                     'RP': ['Fp2', 'F4', 'C4', 'P4', 'O2']}\n\n    # Filter specifications\n    nyquist_freq = 0.5 * 200\n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    # Bandpass and notch filter\n    # bandpass_coefficients = butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized], btype='band')\n    notch_coefficients = iirnotch(w0=60, Q=30, fs=200)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype='band')\n\n    spec_size = duration * 200\n    start = start * 200\n    real_start = start + (10_000 // 2) - (spec_size // 2)\n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    # Spectrogram parameters\n    fs = 200\n\n    if spec_size_freq <= 0 or spec_size_time <= 0:\n        freq_size = int((nfft // 2) / 5.15198) + 1\n        segments = int((spec_size - noverlap) / (nperseg - noverlap))\n    else:\n        freq_size = spec_size_freq\n        segments = spec_size_time\n\n    # Initialize spectrogram container\n    spectrogram = cp.zeros((freq_size, segments, 4), dtype='float32')\n\n    processed_eeg = {}\n\n    for i, (electrode_pair_name, electrode_locs) in enumerate(electrode_pair_name_locations.items()):\n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n        for j in range(4):\n            # Compute differential signals\n            signal = cp.array(eeg_data[electrode_locs[j]].values - eeg_data[electrode_locs[j + 1]].values)\n\n            # Handles NaNs \n            mean_signal = cp.nanmean(signal)\n            signal = cp.nan_to_num(signal, nan=mean_signal) if cp.isnan(signal).mean() < 1 else cp.zeros_like(signal)\n\n            # Filters bandpass and notch\n            signal_filtered = filtfilt(*notch_coefficients, signal)\n            signal_filtered = scipy_filtfilt(*sci_bandpass_coefficients, signal_filtered.get())  # HOTFIX\n\n            # GPU-accelerated spectrogram computation\n            frequencies, times, Sxx = cupyx_spectrogram(signal_filtered, fs, nperseg=nperseg, noverlap=noverlap,\n                                                        nfft=nfft)\n\n            # Filters frequency range \n            valid_freq = (frequencies >= 0.59) & (frequencies <= 20)\n            Sxx_filtered = Sxx[valid_freq, :]\n\n            # Logarithmic transformation and normalization using Cupy\n            spectrogram_slice = cp.clip(Sxx_filtered, cp.exp(-4), cp.exp(6))\n            spectrogram_slice = cp.log10(spectrogram_slice)\n\n            normalization_epsilon = 1e-6\n            mean = spectrogram_slice.mean(axis=(0, 1), keepdims=True)\n            std = spectrogram_slice.std(axis=(0, 1), keepdims=True)\n            spectrogram_slice = (spectrogram_slice - mean) / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[f'{electrode_locs[j]}_{electrode_locs[j + 1]}'] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        # AVERAGES THE 4 MONTAGE DIFFERENCES\n        if mean_montage_names > 0:\n            spectrogram[:, :, i] /= mean_montage_names\n\n    # Applies Gaussian filter and retrieves the spectrogram as a NumPy array using cupy.ndarray.get()\n    spec_numpy = gaussian_filter(spectrogram, sigma=sigma_gaussian).get() if sigma_gaussian > 0 else spectrogram.get()\n\n    # Filter EKG signal\n    ekg_signal_filtered = filtfilt(*notch_coefficients, cp.array(eeg_data[\"EKG\"].values))\n    processed_eeg['EKG'] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered.get())  # HOTFIX\n    return spec_numpy, processed_eeg\n```\n\n##Bonus\nAlso, after preparing 50 sec and 10 sec specs you can create the final image differently. So for example, rather than resizing it you can cut high frequencies like shown below and center left-right parts mirror-like.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Feefd98c6ccc8b8a1afb782d0952175fe%2F1.png?generation=1711557383664757&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ffeb52985983b67376f4aeef64742937e%2F2.png?generation=1711557395408198&alt=media)\n\nThough the originally proposed method used by Rafael brings the best score for us. You are welcome to share your ideas on how you combine the thing, but I am pretty much sure top teams did not leave Rafael's notebook unnoticed.\n\n##Outro\nHave fun and good luck, there is still a time to improve!",
    "2719358": "P.s. Let me know if you want me to post the rewritten Rafael's notebook for data preparation with both head cropping part and his original method.",
    "2735113": "Hey friends, I made a bug fix in the model that resulted in an LB of 0.31. Below is the link to the discussion where I explain what happened :)\n\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/491070",
    "2747734": "@sergiosaharovskiy cusignal is deprecated by NVIDIA. The RAPIDS team has contributed its content to cupy.\n\nThis is to say that you are absolutely right! One should now use cupy to perform what cusignal was doing. NVIDIA developed cusignal to show that it is possible to perform signal processing on GPU. Moving it to cupy was the right way to make it both simpler to use, and more accessible.",
    "2734973": "Thanks for the implementation. It does indeed work faster. But I was hoping that your and Rafael's approach would give a result close to 0.32lb. Hmm, maybe I'm doing something wrong)",
    "2722528": "Hi Sergey, may I ask which version of cupy are you installing? I tried multiple versions, either some of the following  are missing \n\nfrom cupyx.scipy.ndimage import gaussian_filter\nfrom cupyx.scipy.signal import filtfilt, iirnotch\nfrom cupyx.scipy.signal import spectrogram as cupyx_spectrogram\n\nor requiring some `lbnvrtc.so` stuff, very annoying...\n\nThank you!",
    "2721690": "Thanks for sharing!\nDid you do use the same preprocessing’s parameter with the Rafael’s notebook?",
    "2720985": "Thanks for sharing! Did you do any preprocessing to `eeg_data` before passing it into the function? With `eeg_data` being `pd.read_parquet(...)`, and calling your function, I get the error: `ValueError: operands could not be broadcast together with shapes (267, 30) (267, 501) (267, 30)` on line `spectrogram[i, :, :] += spectrogram_slice`.\n\nNot sure if you've trimmed the loaded `eeg_data` before calling that function?",
    "2719621": "Thank you for the mention, Sergey. I thought your code optimization was quite good. Indeed, installing Rapids can sometimes be slow and end up being a hassle, but this way, it's much easier to use.",
    "2736821": "Throw multiple error with me outside KAGGLE, the fix from my point of view is below, however it may be experience low performance or different results.\n\n```\ndef create_spectrogram_with_cupy(eeg_data, start, duration=50,\n                                 low_cut_freq=0.7, high_cut_freq=20, order_band=5,\n                                 spec_size_freq=267, spec_size_time=30,\n                                 nperseg=1500, noverlap=1483, nfft=2750,\n                                 sigma_gaussian=0.7,\n                                 mean_montage_names=4):\n    electrode_pair_name_locations = {'LL': ['Fp1', 'F7', 'T3', 'T5', 'O1'],\n                                     'RL': ['Fp2', 'F8', 'T4', 'T6', 'O2'],\n                                     'LP': ['Fp1', 'F3', 'C3', 'P3', 'O1'],\n                                     'RP': ['Fp2', 'F4', 'C4', 'P4', 'O2']}\n\n    # Filter specifications\n    nyquist_freq = 0.5 * 200\n    low_cut_freq_normalized = low_cut_freq / nyquist_freq\n    high_cut_freq_normalized = high_cut_freq / nyquist_freq\n\n    # Bandpass and notch filter\n    # bandpass_coefficients = butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized], btype='band')\n    notch_coefficients = iirnotch(w0=60, Q=30, fs=200)\n    sci_bandpass_coefficients = scipy_butter(order_band, [low_cut_freq_normalized, high_cut_freq_normalized],\n                                             btype='band')\n\n    spec_size = duration * 200\n    start = start * 200\n    real_start = start + (10_000 // 2) - (spec_size // 2)\n    eeg_data = eeg_data.iloc[real_start:real_start + spec_size]\n\n    # Spectrogram parameters\n    fs = 200\n\n    if spec_size_freq <= 0 or spec_size_time <= 0:\n        freq_size = int((nfft // 2) / 5.15198) + 1\n        segments = int((spec_size - noverlap) / (nperseg - noverlap))\n    else:\n        freq_size = spec_size_freq\n        segments = spec_size_time\n\n    # Initialize spectrogram container\n    spectrogram = cp.zeros((freq_size, segments, 4), dtype='float32')\n\n    processed_eeg = {}\n\n    for i, (electrode_pair_name, electrode_locs) in enumerate(electrode_pair_name_locations.items()):\n        processed_eeg[electrode_pair_name] = np.zeros(spec_size)\n\n        for j in range(4):\n            # Compute differential signals\n            signal = cp.array(eeg_data[electrode_locs[j]].values - eeg_data[electrode_locs[j + 1]].values)\n\n            # Handles NaNs \n            mean_signal = cp.nanmean(signal)\n            signal = cp.nan_to_num(signal, nan=mean_signal) if cp.isnan(signal).mean() < 1 else cp.zeros_like(signal)\n\n            signal_np = signal.get()\n\n            # Filters bandpass and notch\n            signal_filtered = filtfilt(*notch_coefficients, signal_np)\n            signal_filtered = scipy_filtfilt(*sci_bandpass_coefficients, signal_filtered)  # HOTFIX\n\n            # GPU-accelerated spectrogram computation\n            frequencies, times, Sxx = cupyx_spectrogram(signal_filtered, fs, nperseg=nperseg, noverlap=noverlap,\n                                                        nfft=nfft)\n\n            # Filters frequency range \n            valid_freq = (frequencies >= 0.59) & (frequencies <= 20)\n            Sxx_filtered = Sxx[valid_freq, :]\n\n            # Logarithmic transformation and normalization using Cupy\n            spectrogram_slice = cp.clip(Sxx_filtered, cp.exp(-4), cp.exp(6))\n            spectrogram_slice = cp.log10(spectrogram_slice)\n\n            normalization_epsilon = 1e-6\n            mean = spectrogram_slice.mean(axis=(0, 1), keepdims=True)\n            std = spectrogram_slice.std(axis=(0, 1), keepdims=True)\n            spectrogram_slice = (spectrogram_slice - mean) / (std + normalization_epsilon)\n\n            spectrogram[:, :, i] += spectrogram_slice\n            processed_eeg[f'{electrode_locs[j]}_{electrode_locs[j + 1]}'] = signal.get()\n            processed_eeg[electrode_pair_name] += signal.get()\n\n        # AVERAGES THE 4 MONTAGE DIFFERENCES\n        if mean_montage_names > 0:\n            spectrogram[:, :, i] /= mean_montage_names\n\n    # Applies Gaussian filter and retrieves the spectrogram as a NumPy array using cupy.ndarray.get()\n    spec_numpy = gaussian_filter(spectrogram, sigma=sigma_gaussian).get() if sigma_gaussian > 0 else spectrogram.get()\n\n    # Filter EKG signal\n    ekg_np = cp.array(eeg_data[\"EKG\"].values).get()\n    ekg_signal_filtered = filtfilt(*notch_coefficients, ekg_np )\n    processed_eeg['EKG'] = scipy_filtfilt(*sci_bandpass_coefficients, ekg_signal_filtered)  # HOTFIX\n    return spec_numpy, processed_eeg\n\n```",
    "2720308": "Thanks for sharing.\nMy experiment is also going well with Rafael's data.\nDo you use masking augmentation ? \n ",
    "2723327": "Hi Qui,\n\nLocally I use the following:\n\ncupy-cuda11x==13.0.0\nnvidia-cuda-nvrtc-cu11==11.7.99\nnvidia-cuda-cupti-cu11==11.7.101\nnvidia-cuda-runtime-cu11==11.7.99",
    "2720795": "@sergiosaharovskiy with Melspectrogram I know I can restrict output spectrogram to a certain size, but cupy signal spectrogram seems to be a little tricky. Is there any way to control the output size of the spectrogram?",
    "2719895": "",
    "2723867": "thanks for sharing!"
  }
}