{
  "id": 473077,
  "title": "Kaggle spectrogram from spectrogram parquet file",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/473077",
  "author_name": "",
  "post_date": "2024-02-03T11:10:32.144673200Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>How do you construct the spectrogram image from spectrogram parquet file? And also if we construct spectrogram from EEG we use 50 seconds of data. But in test data, the kaggle spectrogram use the 10 min.</p>",
  "messages": [
    {
      "id": "2633909",
      "postDate": "02/03/2024 11:10:32",
      "content": "<p>How do you construct the spectrogram image from spectrogram parquet file? And also if we construct spectrogram from EEG we use 50 seconds of data. But in test data, the kaggle spectrogram use the 10 min.</p>",
      "rawMarkdown": "How do you construct the spectrogram image from spectrogram parquet file? And also if we construct spectrogram from EEG we use 50 seconds of data. But in test data, the kaggle spectrogram use the 10 min.",
      "votes": null
    },
    {
      "id": "2634311",
      "postDate": "02/03/2024 17:15:48",
      "content": "<p><a href=\"https://www.kaggle.com/pritamsinha23\" target=\"_blank\">@pritamsinha23</a> <br>\n<br>\n</p>\n<ul>\n<li>i miss understood the question, Chris answered.</li>\n</ul>",
      "rawMarkdown": "pritamsinha23 \n~~1. 10-minute spectrograms impossible due to incomplete EEG data.~~\n~~2. Creating spectrograms from 50-second sub sequence is possible.~~\n- i miss understood the question, Chris answered.",
      "votes": null
    },
    {
      "id": "2634324",
      "postDate": "02/03/2024 17:31:03",
      "content": "<p>The image is just <code>img = df.iloc[start:start+300,:].values.T</code> where <code>df</code> is Kaggle's spectrogram parquet. Afterward the variable <code>img</code> is a NumPy array of dimension <code>(400x300)</code> which is <code>(frequency x time)</code>. The frequencies range from 0 to 20Hz. And the time ranges from 0 to 10 minutes.</p>\n<p>We can then plot with <code>plt.imshow(img)</code>. Or we can make it look better with log transform: </p>\n<pre><code> = np(,np(-),np())\n = np(img)\nplt(img)\n</code></pre>",
      "rawMarkdown": "The image is just `img = df.iloc[start:start+300,:].values.T` where `df` is Kaggle's spectrogram parquet. Afterward the variable `img` is a NumPy array of dimension `(400x300)` which is `(frequency x time)`. The frequencies range from 0 to 20Hz. And the time ranges from 0 to 10 minutes.\n\nWe can then plot with `plt.imshow(img)`. Or we can make it look better with log transform: \n\n    img = np.clip(img,np.exp(-4),np.exp(8))\n    img = np.log(img)\n    plt.imshow(img)",
      "votes": null
    },
    {
      "id": "2634484",
      "postDate": "02/03/2024 19:17:05",
      "content": "<p>Thanks Chris for the reply. One question, I am assuming the <code>start</code> variable is the <code>spectrogram_offset_labels_seconds</code>. And why <code>300</code>. How is it equivalent to <code>10 minutes</code>? Can you clarify this please.</p>",
      "rawMarkdown": "Thanks Chris for the reply. One question, I am assuming the `start` variable is the `spectrogram_offset_labels_seconds`. And why `300`. How is it equivalent to `10 minutes`? Can you clarify this please.",
      "votes": null
    },
    {
      "id": "2634494",
      "postDate": "02/03/2024 19:22:23",
      "content": "<p>We use <code>300</code> because each row of spectrogram parquet is 2 seconds. Therefore <code>300 x 2 seconds / 60 = 10 minutes</code>. The start is <code>spectrogram_offset_labels_seconds</code>. But we need to include the factor of 2 when converting to row number. The safest way to get the correct rows is to use Pandas location with</p>\n<pre><code>df.loc[(df.&gt;=spectrogram_offset_labels_seconds)&amp;(df.&lt;spectrogram_offset_labels_seconds+)]\n</code></pre>\n<p>That will extract the correct 300 rows.</p>",
      "rawMarkdown": "We use `300` because each row of spectrogram parquet is 2 seconds. Therefore `300 x 2 seconds / 60 = 10 minutes`. The start is `spectrogram_offset_labels_seconds`. But we need to include the factor of 2 when converting to row number. The safest way to get the correct rows is to use Pandas location with\n\n    df.loc[(df.time>=spectrogram_offset_labels_seconds)&(df.time<spectrogram_offset_labels_seconds+600)]\n\nThat will extract the correct 300 rows.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2634311,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "02/03/2024 17:15:48",
      "content": "<p><a href=\"https://www.kaggle.com/pritamsinha23\" target=\"_blank\">@pritamsinha23</a> <br>\n<br>\n</p>\n<ul>\n<li>i miss understood the question, Chris answered.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2634324,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/03/2024 17:31:03",
      "content": "<p>The image is just <code>img = df.iloc[start:start+300,:].values.T</code> where <code>df</code> is Kaggle's spectrogram parquet. Afterward the variable <code>img</code> is a NumPy array of dimension <code>(400x300)</code> which is <code>(frequency x time)</code>. The frequencies range from 0 to 20Hz. And the time ranges from 0 to 10 minutes.</p>\n<p>We can then plot with <code>plt.imshow(img)</code>. Or we can make it look better with log transform: </p>\n<pre><code> = np(,np(-),np())\n = np(img)\nplt(img)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2634484,
          "author_name": "pritamsinha23",
          "author_url": "",
          "post_date": "02/03/2024 19:17:05",
          "content": "<p>Thanks Chris for the reply. One question, I am assuming the <code>start</code> variable is the <code>spectrogram_offset_labels_seconds</code>. And why <code>300</code>. How is it equivalent to <code>10 minutes</code>? Can you clarify this please.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2634494,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/03/2024 19:22:23",
              "content": "<p>We use <code>300</code> because each row of spectrogram parquet is 2 seconds. Therefore <code>300 x 2 seconds / 60 = 10 minutes</code>. The start is <code>spectrogram_offset_labels_seconds</code>. But we need to include the factor of 2 when converting to row number. The safest way to get the correct rows is to use Pandas location with</p>\n<pre><code>df.loc[(df.&gt;=spectrogram_offset_labels_seconds)&amp;(df.&lt;spectrogram_offset_labels_seconds+)]\n</code></pre>\n<p>That will extract the correct 300 rows.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2633909": "How do you construct the spectrogram image from spectrogram parquet file? And also if we construct spectrogram from EEG we use 50 seconds of data. But in test data, the kaggle spectrogram use the 10 min.",
    "2634311": "pritamsinha23 \n~~1. 10-minute spectrograms impossible due to incomplete EEG data.~~\n~~2. Creating spectrograms from 50-second sub sequence is possible.~~\n- i miss understood the question, Chris answered.",
    "2634324": "The image is just `img = df.iloc[start:start+300,:].values.T` where `df` is Kaggle's spectrogram parquet. Afterward the variable `img` is a NumPy array of dimension `(400x300)` which is `(frequency x time)`. The frequencies range from 0 to 20Hz. And the time ranges from 0 to 10 minutes.\n\nWe can then plot with `plt.imshow(img)`. Or we can make it look better with log transform: \n\n    img = np.clip(img,np.exp(-4),np.exp(8))\n    img = np.log(img)\n    plt.imshow(img)",
    "2634484": "Thanks Chris for the reply. One question, I am assuming the `start` variable is the `spectrogram_offset_labels_seconds`. And why `300`. How is it equivalent to `10 minutes`? Can you clarify this please.",
    "2634494": "We use `300` because each row of spectrogram parquet is 2 seconds. Therefore `300 x 2 seconds / 60 = 10 minutes`. The start is `spectrogram_offset_labels_seconds`. But we need to include the factor of 2 when converting to row number. The safest way to get the correct rows is to use Pandas location with\n\n    df.loc[(df.time>=spectrogram_offset_labels_seconds)&(df.time<spectrogram_offset_labels_seconds+600)]\n\nThat will extract the correct 300 rows."
  },
  "source": "meta"
}