{
  "id": 484794,
  "title": "Spectrogram processing",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/484794",
  "author_name": "",
  "post_date": "2024-03-18T08:48:54.284756100Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I am trying to understand what the signals in spectrogram files are, but when I plot them they appear quite different from results of other people? Is there a correct format on processing them?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4307169%2F825c770925cb8db2c33e09c75714a764%2Fspectrogram%20HMS.png?generation=1710751661575285&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2703523",
      "postDate": "03/18/2024 08:48:54",
      "content": "<p>I am trying to understand what the signals in spectrogram files are, but when I plot them they appear quite different from results of other people? Is there a correct format on processing them?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4307169%2F825c770925cb8db2c33e09c75714a764%2Fspectrogram%20HMS.png?generation=1710751661575285&amp;alt=media\"></p>",
      "rawMarkdown": "I am trying to understand what the signals in spectrogram files are, but when I plot them they appear quite different from results of other people? Is there a correct format on processing them?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4307169%2F825c770925cb8db2c33e09c75714a764%2Fspectrogram%20HMS.png?generation=1710751661575285&alt=media)",
      "votes": null
    },
    {
      "id": "2703574",
      "postDate": "03/18/2024 09:05:15",
      "content": "<p>You should do log transform.</p>",
      "rawMarkdown": "You should do log transform.",
      "votes": null
    },
    {
      "id": "2703591",
      "postDate": "03/18/2024 09:12:44",
      "content": "<p>it encountered problem with dividing by zero, what value should I replace zeros with?</p>",
      "rawMarkdown": "it encountered problem with dividing by zero, what value should I replace zeros with?",
      "votes": null
    },
    {
      "id": "2703605",
      "postDate": "03/18/2024 09:30:26",
      "content": "<p>You need a cutoff for the log transform since log(x-&gt;0)-&gt;-infinity. For example log(np.clip(x, a_min = cutoff)). Alternatively, you can just log(x+epsilon) assuming that x&gt;=0 (should hold for spectrograms). Try for example with cutoff/epsilon = 10^(-10).</p>",
      "rawMarkdown": "You need a cutoff for the log transform since log(x->0)->-infinity. For example log(np.clip(x, a_min = cutoff)). Alternatively, you can just log(x+epsilon) assuming that x>=0 (should hold for spectrograms). Try for example with cutoff/epsilon = 10^(-10).",
      "votes": null
    },
    {
      "id": "2703612",
      "postDate": "03/18/2024 09:38:09",
      "content": "<p>try <code>np.log1p(x)</code> to avoid log(0).</p>",
      "rawMarkdown": "try `np.log1p(x)` to avoid log(0).",
      "votes": null
    },
    {
      "id": "2704235",
      "postDate": "03/18/2024 16:56:12",
      "content": "<p>Three simple steps. Given a spectrogram \"S\":</p>\n<p>1st - Clip the values at 0 &lt; S &lt; +Inf, this will make sense in the next step;</p>\n<p>2nd - Do a log transform to convert the values from Watts to dB: do 10*log10(S/max(S)). You can use any log base, the only difference between bases is the scale factor, but log10 is the most common. This will kind of normalize the values and reduce the large values differences from the original spectrogram. log(x) is undefined at 0, which is why step #1 is necessary;</p>\n<p>3rd (optional) - Clip the spectrogram at a top value (top_dB): max(10*log10(S/max(S)))-top_dB.</p>\n<p>If you follow all 3 of these steps your spectrogram's values will be in the range[-top_dB, 0]. That way you can normalize them more easily (e.g. if you want the values of your spectrogram to be in the [0,1] range simply divide the spectrogram by -top_dB). Keep in mind that each Kaggle spectrogram is actually 4 different spectrograms, so you should process them individually.</p>",
      "rawMarkdown": "Three simple steps. Given a spectrogram \"S\":\n\n1st - Clip the values at 0 < S < +Inf, this will make sense in the next step;\n\n2nd - Do a log transform to convert the values from Watts to dB: do 10*log10(S/max(S)). You can use any log base, the only difference between bases is the scale factor, but log10 is the most common. This will kind of normalize the values and reduce the large values differences from the original spectrogram. log(x) is undefined at 0, which is why step #1 is necessary;\n\n3rd (optional) - Clip the spectrogram at a top value (top_dB): max(10*log10(S/max(S)))-top_dB.\n\nIf you follow all 3 of these steps your spectrogram's values will be in the range[-top_dB, 0]. That way you can normalize them more easily (e.g. if you want the values of your spectrogram to be in the [0,1] range simply divide the spectrogram by -top_dB). Keep in mind that each Kaggle spectrogram is actually 4 different spectrograms, so you should process them individually.",
      "votes": null
    },
    {
      "id": "2705508",
      "postDate": "03/19/2024 11:43:10",
      "content": "<p>thank you for the explanation, I will implement it and see what happens</p>",
      "rawMarkdown": "thank you for the explanation, I will implement it and see what happens",
      "votes": null
    },
    {
      "id": "2705633",
      "postDate": "03/19/2024 12:32:58",
      "content": "<p>I want to ask, what's top_dB? what value is that? where do I get it from?</p>",
      "rawMarkdown": "I want to ask, what's top_dB? what value is that? where do I get it from?",
      "votes": null
    },
    {
      "id": "2706171",
      "postDate": "03/19/2024 17:45:41",
      "content": "<p>top_dB is an arbitrary value, it is meant to threshold the top value and it also makes it easier to normalize the values before feeding the data to a model. The <a href=\"https://librosa.org/\" target=\"_blank\">librosa</a> library uses top_dB = 80, but you could go higher (or lower) if you want to, like 96 or 64. I forgot to link the <a href=\"https://librosa.org/doc/main/generated/librosa.power_to_db.html\" target=\"_blank\">librosa's function</a> in my original post, but I think that creating the function yourself will give you more flexibility if you are not working with a 2D array (multiple spectrograms per array or an <a href=\"https://mne.tools/dev/index.html\" target=\"_blank\">MNE</a> spectrogram).</p>",
      "rawMarkdown": "top_dB is an arbitrary value, it is meant to threshold the top value and it also makes it easier to normalize the values before feeding the data to a model. The [librosa](https://librosa.org/) library uses top_dB = 80, but you could go higher (or lower) if you want to, like 96 or 64. I forgot to link the [librosa's function](https://librosa.org/doc/main/generated/librosa.power_to_db.html) in my original post, but I think that creating the function yourself will give you more flexibility if you are not working with a 2D array (multiple spectrograms per array or an [MNE](https://mne.tools/dev/index.html) spectrogram).",
      "votes": null
    },
    {
      "id": "2706241",
      "postDate": "03/19/2024 18:47:19",
      "content": "<p>Got it, thank you so much!</p>",
      "rawMarkdown": "Got it, thank you so much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2703574,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "03/18/2024 09:05:15",
      "content": "<p>You should do log transform.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2703591,
          "author_name": "anrenk",
          "author_url": "",
          "post_date": "03/18/2024 09:12:44",
          "content": "<p>it encountered problem with dividing by zero, what value should I replace zeros with?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2703605,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "03/18/2024 09:30:26",
              "content": "<p>You need a cutoff for the log transform since log(x-&gt;0)-&gt;-infinity. For example log(np.clip(x, a_min = cutoff)). Alternatively, you can just log(x+epsilon) assuming that x&gt;=0 (should hold for spectrograms). Try for example with cutoff/epsilon = 10^(-10).</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2703612,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/18/2024 09:38:09",
      "content": "<p>try <code>np.log1p(x)</code> to avoid log(0).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2704235,
      "author_name": "gabrielribeirocsilva",
      "author_url": "",
      "post_date": "03/18/2024 16:56:12",
      "content": "<p>Three simple steps. Given a spectrogram \"S\":</p>\n<p>1st - Clip the values at 0 &lt; S &lt; +Inf, this will make sense in the next step;</p>\n<p>2nd - Do a log transform to convert the values from Watts to dB: do 10*log10(S/max(S)). You can use any log base, the only difference between bases is the scale factor, but log10 is the most common. This will kind of normalize the values and reduce the large values differences from the original spectrogram. log(x) is undefined at 0, which is why step #1 is necessary;</p>\n<p>3rd (optional) - Clip the spectrogram at a top value (top_dB): max(10*log10(S/max(S)))-top_dB.</p>\n<p>If you follow all 3 of these steps your spectrogram's values will be in the range[-top_dB, 0]. That way you can normalize them more easily (e.g. if you want the values of your spectrogram to be in the [0,1] range simply divide the spectrogram by -top_dB). Keep in mind that each Kaggle spectrogram is actually 4 different spectrograms, so you should process them individually.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2705508,
          "author_name": "anrenk",
          "author_url": "",
          "post_date": "03/19/2024 11:43:10",
          "content": "<p>thank you for the explanation, I will implement it and see what happens</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2705633,
          "author_name": "anrenk",
          "author_url": "",
          "post_date": "03/19/2024 12:32:58",
          "content": "<p>I want to ask, what's top_dB? what value is that? where do I get it from?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2706171,
              "author_name": "gabrielribeirocsilva",
              "author_url": "",
              "post_date": "03/19/2024 17:45:41",
              "content": "<p>top_dB is an arbitrary value, it is meant to threshold the top value and it also makes it easier to normalize the values before feeding the data to a model. The <a href=\"https://librosa.org/\" target=\"_blank\">librosa</a> library uses top_dB = 80, but you could go higher (or lower) if you want to, like 96 or 64. I forgot to link the <a href=\"https://librosa.org/doc/main/generated/librosa.power_to_db.html\" target=\"_blank\">librosa's function</a> in my original post, but I think that creating the function yourself will give you more flexibility if you are not working with a 2D array (multiple spectrograms per array or an <a href=\"https://mne.tools/dev/index.html\" target=\"_blank\">MNE</a> spectrogram).</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2706241,
                  "author_name": "anrenk",
                  "author_url": "",
                  "post_date": "03/19/2024 18:47:19",
                  "content": "<p>Got it, thank you so much!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2703523": "I am trying to understand what the signals in spectrogram files are, but when I plot them they appear quite different from results of other people? Is there a correct format on processing them?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4307169%2F825c770925cb8db2c33e09c75714a764%2Fspectrogram%20HMS.png?generation=1710751661575285&alt=media)",
    "2703574": "You should do log transform.",
    "2703591": "it encountered problem with dividing by zero, what value should I replace zeros with?",
    "2703605": "You need a cutoff for the log transform since log(x->0)->-infinity. For example log(np.clip(x, a_min = cutoff)). Alternatively, you can just log(x+epsilon) assuming that x>=0 (should hold for spectrograms). Try for example with cutoff/epsilon = 10^(-10).",
    "2703612": "try `np.log1p(x)` to avoid log(0).",
    "2704235": "Three simple steps. Given a spectrogram \"S\":\n\n1st - Clip the values at 0 < S < +Inf, this will make sense in the next step;\n\n2nd - Do a log transform to convert the values from Watts to dB: do 10*log10(S/max(S)). You can use any log base, the only difference between bases is the scale factor, but log10 is the most common. This will kind of normalize the values and reduce the large values differences from the original spectrogram. log(x) is undefined at 0, which is why step #1 is necessary;\n\n3rd (optional) - Clip the spectrogram at a top value (top_dB): max(10*log10(S/max(S)))-top_dB.\n\nIf you follow all 3 of these steps your spectrogram's values will be in the range[-top_dB, 0]. That way you can normalize them more easily (e.g. if you want the values of your spectrogram to be in the [0,1] range simply divide the spectrogram by -top_dB). Keep in mind that each Kaggle spectrogram is actually 4 different spectrograms, so you should process them individually.",
    "2705508": "thank you for the explanation, I will implement it and see what happens",
    "2705633": "I want to ask, what's top_dB? what value is that? where do I get it from?",
    "2706171": "top_dB is an arbitrary value, it is meant to threshold the top value and it also makes it easier to normalize the values before feeding the data to a model. The [librosa](https://librosa.org/) library uses top_dB = 80, but you could go higher (or lower) if you want to, like 96 or 64. I forgot to link the [librosa's function](https://librosa.org/doc/main/generated/librosa.power_to_db.html) in my original post, but I think that creating the function yourself will give you more flexibility if you are not working with a 2D array (multiple spectrograms per array or an [MNE](https://mne.tools/dev/index.html) spectrogram).",
    "2706241": "Got it, thank you so much!"
  },
  "source": "meta"
}