{
  "id": 230865,
  "title": "Pre-computed Spectrograms Dataset",
  "url": "/competitions/birdclef-2021/discussion/230865",
  "author_name": "Takamichi Toda",
  "post_date": "2021-04-06T00:18:38.629000",
  "votes": 36,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I think this competition needs lots of computing time to calculate the spectrograms.<br>\nWhen we pre-compute the spectrograms, we can save computing time.</p>\n<p>I publish pre-computed spectrograms dataset.  <br>\nThe dataset was very large, I split it by the initial alphabet.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-01-a\" target=\"_blank\">BirdCLEF Spectrogram 01 (a)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-02-b\" target=\"_blank\">BirdCLEF Spectrogram 02 (b)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-03-c\" target=\"_blank\">BirdCLEF Spectrogram 03 (c)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-04-dg\" target=\"_blank\">BirdCLEF Spectrogram 04 (d-g)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-05-hm\" target=\"_blank\">BirdCLEF Spectrogram 05 (h-m)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-06-np\" target=\"_blank\">BirdCLEF Spectrogram 06 (n-p)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-07-qs\" target=\"_blank\">BirdCLEF Spectrogram 07 (q-s)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-08-tz\" target=\"_blank\">BirdCLEF Spectrogram 08 (t-z)</a></li>\n</ul>\n<p>I calculate spectrograms by using this code.</p>\n<pre><code>class config:\n    INPUT_ROOT = \"/kaggle/input/birdclef-2021\"\n    WORK_ROOT = \"/kaggle/working\"\n    FMIN = 20\n    FMAX = 16000\n    SPEC_HEIGHT = 128\n    N_FFT = 2048\n</code></pre>\n<pre><code>def load_mel_spec(path):\n    data, samplerate = sf.read(path)\n    mel_spec = librosa.feature.melspectrogram(y=data, \n                                              sr=samplerate, \n                                              n_fft=config.N_FFT, \n                                              n_mels=config.SPEC_HEIGHT, \n                                              fmin=config.FMIN, \n                                              fmax=config.FMAX)\n    mel_spec = librosa.power_to_db(mel_spec, ref=np.max) \n    return mel_spec\n</code></pre>\n<pre><code>def save_npz(primary_label, filename):\n    path = f\"{config.INPUT_ROOT}/train_short_audio/{primary_label}/{filename}\"\n    mel_spec = load_mel_spec(path)\n    np.save(f\"{config.WORK_ROOT}/{primary_label}/{filename}\", mel_spec)\n</code></pre>\n<p>I will make the starter code with this dataset.</p>",
  "messages": [
    {
      "id": 1264163,
      "postDate": "2021-04-06T00:18:38.630Z",
      "content": "<p>I think this competition needs lots of computing time to calculate the spectrograms.<br>\nWhen we pre-compute the spectrograms, we can save computing time.</p>\n<p>I publish pre-computed spectrograms dataset.  <br>\nThe dataset was very large, I split it by the initial alphabet.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-01-a\" target=\"_blank\">BirdCLEF Spectrogram 01 (a)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-02-b\" target=\"_blank\">BirdCLEF Spectrogram 02 (b)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-03-c\" target=\"_blank\">BirdCLEF Spectrogram 03 (c)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-04-dg\" target=\"_blank\">BirdCLEF Spectrogram 04 (d-g)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-05-hm\" target=\"_blank\">BirdCLEF Spectrogram 05 (h-m)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-06-np\" target=\"_blank\">BirdCLEF Spectrogram 06 (n-p)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-07-qs\" target=\"_blank\">BirdCLEF Spectrogram 07 (q-s)</a></li>\n<li><a href=\"https://www.kaggle.com/takamichitoda/birdclef-spectrogram-08-tz\" target=\"_blank\">BirdCLEF Spectrogram 08 (t-z)</a></li>\n</ul>\n<p>I calculate spectrograms by using this code.</p>\n<pre><code>class config:\n    INPUT_ROOT = \"/kaggle/input/birdclef-2021\"\n    WORK_ROOT = \"/kaggle/working\"\n    FMIN = 20\n    FMAX = 16000\n    SPEC_HEIGHT = 128\n    N_FFT = 2048\n</code></pre>\n<pre><code>def load_mel_spec(path):\n    data, samplerate = sf.read(path)\n    mel_spec = librosa.feature.melspectrogram(y=data, \n                                              sr=samplerate, \n                                              n_fft=config.N_FFT, \n                                              n_mels=config.SPEC_HEIGHT, \n                                              fmin=config.FMIN, \n                                              fmax=config.FMAX)\n    mel_spec = librosa.power_to_db(mel_spec, ref=np.max) \n    return mel_spec\n</code></pre>\n<pre><code>def save_npz(primary_label, filename):\n    path = f\"{config.INPUT_ROOT}/train_short_audio/{primary_label}/{filename}\"\n    mel_spec = load_mel_spec(path)\n    np.save(f\"{config.WORK_ROOT}/{primary_label}/{filename}\", mel_spec)\n</code></pre>\n<p>I will make the starter code with this dataset.</p>",
      "rawMarkdown": "I think this competition needs lots of computing time to calculate the spectrograms.\nWhen we pre-compute the spectrograms, we can save computing time.\n  \nI publish pre-computed spectrograms dataset.  \nThe dataset was very large, I split it by the initial alphabet.\n\n- [BirdCLEF Spectrogram 01 (a)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-01-a)\n- [BirdCLEF Spectrogram 02 (b)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-02-b)\n- [BirdCLEF Spectrogram 03 (c)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-03-c)\n- [BirdCLEF Spectrogram 04 (d-g)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-04-dg)\n- [BirdCLEF Spectrogram 05 (h-m)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-05-hm)\n- [BirdCLEF Spectrogram 06 (n-p)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-06-np)\n- [BirdCLEF Spectrogram 07 (q-s)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-07-qs)\n- [BirdCLEF Spectrogram 08 (t-z)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-08-tz)\n\n  \nI calculate spectrograms by using this code.\n\n```python\nclass config:\n    INPUT_ROOT = \"/kaggle/input/birdclef-2021\"\n    WORK_ROOT = \"/kaggle/working\"\n    FMIN = 20\n    FMAX = 16000\n    SPEC_HEIGHT = 128\n    N_FFT = 2048\n```\n```python\ndef load_mel_spec(path):\n    data, samplerate = sf.read(path)\n    mel_spec = librosa.feature.melspectrogram(y=data, \n                                              sr=samplerate, \n                                              n_fft=config.N_FFT, \n                                              n_mels=config.SPEC_HEIGHT, \n                                              fmin=config.FMIN, \n                                              fmax=config.FMAX)\n    mel_spec = librosa.power_to_db(mel_spec, ref=np.max) \n    return mel_spec\n```\n```python\ndef save_npz(primary_label, filename):\n    path = f\"{config.INPUT_ROOT}/train_short_audio/{primary_label}/{filename}\"\n    mel_spec = load_mel_spec(path)\n    np.save(f\"{config.WORK_ROOT}/{primary_label}/{filename}\", mel_spec)\n```\n\n\nI will make the starter code with this dataset.\n\n",
      "votes": 35
    },
    {
      "id": 1271555,
      "postDate": "2021-04-12T17:33:45.013Z",
      "content": "<p>One more slightly stupid question :</p>\n<p>In the test data, we need to predict for chunks of 5s. However, the spectrogram data here is for the whole audio file as far as I understand it ? Wont this be a problem ?</p>",
      "rawMarkdown": "One more slightly stupid question :\n\nIn the test data, we need to predict for chunks of 5s. However, the spectrogram data here is for the whole audio file as far as I understand it ? Wont this be a problem ?",
      "votes": 1,
      "replies": [
        {
          "id": 1271798,
          "postDate": "2021-04-12T23:18:47.640Z",
          "content": "<p>Yes, this spectrogram data here is for the whole audio file.<br>\nSo, you need to crop the spectrogram for 5 seconds in training.</p>",
          "rawMarkdown": "Yes, this spectrogram data here is for the whole audio file.\nSo, you need to crop the spectrogram for 5 seconds in training.\n"
        },
        {
          "id": 1272859,
          "postDate": "2021-04-13T20:26:00.633Z",
          "content": "<p>Just to lightly seed another line of thought:<br>\nWhen you process the soundscapes your model makes predictions on a 5s window, but does have access to the full file. There may be nice ways to take advantage of longer time windows when predicting a 5s window. I don't /know/ that there's an approach that will work, but I do think it's a bit under-explored.</p>",
          "rawMarkdown": "Just to lightly seed another line of thought:\nWhen you process the soundscapes your model makes predictions on a 5s window, but does have access to the full file. There may be nice ways to take advantage of longer time windows when predicting a 5s window. I don't /know/ that there's an approach that will work, but I do think it's a bit under-explored.",
          "votes": 4
        },
        {
          "id": 1273019,
          "postDate": "2021-04-14T02:51:06.780Z",
          "content": "<p>Thank you for your advice!<br>\nI may be able to make some predictions from the audio before and after.<br>\nI'll challenge!</p>",
          "rawMarkdown": "Thank you for your advice!\nI may be able to make some predictions from the audio before and after.\nI'll challenge!"
        },
        {
          "id": 1273454,
          "postDate": "2021-04-14T11:16:09.043Z",
          "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> To your point, the SED model that was heavily used in previous competition and is also reused in some public notebooks here does look at more than 5 second clips.  Don't you agree?</p>",
          "rawMarkdown": "@tomdenton To your point, the SED model that was heavily used in previous competition and is also reused in some public notebooks here does look at more than 5 second clips.  Don't you agree?",
          "votes": 3
        },
        {
          "id": 1274020,
          "postDate": "2021-04-14T22:47:22.393Z",
          "content": "<p>FYI, prediction on 30s window was better than prediction on 5s window for me in the previous competition.</p>",
          "rawMarkdown": "FYI, prediction on 30s window was better than prediction on 5s window for me in the previous competition.",
          "votes": 6
        },
        {
          "id": 1276734,
          "postDate": "2021-04-17T22:36:22.953Z",
          "content": "<p>As a birder, the use of longer time segments makes a lot of intuitive sense. There are a lot of cases in which pattern of repetition matters as much or more than the the sound of a single element or phrase. A single element on its own heard at a distance might be indistinguishable between a Red-eyed Vireo, a Blue-headed Vireo, a Yellow-throated Vireo, a Philadelphia Vireo, or even an American Robin. But if the bird keeps singing, incessantly, particularly if is a summer afternoon, then Red-eyed Vireo is a good bet (assuming you’re in its range, of course, and the habitat is right).  Likewise, a single 5 second chunk of a Northern Mockingbird could be really confusing, but as it keeps singing, the mystery vanishes (at least to a listening birder); mockingbirds often repeat elements a few times before switching to a different pattern for a few repetitions, and on and on. Though the content varies, the pattern is reliable. </p>\n<p>I have to admit I’m still learning my way around the options for implementing this,  but my instinct says it would be worth the effort. Thanks to those who have offered up some leads. </p>",
          "rawMarkdown": "As a birder, the use of longer time segments makes a lot of intuitive sense. There are a lot of cases in which pattern of repetition matters as much or more than the the sound of a single element or phrase. A single element on its own heard at a distance might be indistinguishable between a Red-eyed Vireo, a Blue-headed Vireo, a Yellow-throated Vireo, a Philadelphia Vireo, or even an American Robin. But if the bird keeps singing, incessantly, particularly if is a summer afternoon, then Red-eyed Vireo is a good bet (assuming you’re in its range, of course, and the habitat is right).  Likewise, a single 5 second chunk of a Northern Mockingbird could be really confusing, but as it keeps singing, the mystery vanishes (at least to a listening birder); mockingbirds often repeat elements a few times before switching to a different pattern for a few repetitions, and on and on. Though the content varies, the pattern is reliable. \n\nI have to admit I’m still learning my way around the options for implementing this,  but my instinct says it would be worth the effort. Thanks to those who have offered up some leads. ",
          "votes": 5
        },
        {
          "id": 1320433,
          "postDate": "2021-05-24T04:52:27.560Z",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>  i was not part of previous competition i did thought of using long er duration but problem is how we break predictions  and assign them to 5 sec clips.</p>",
          "rawMarkdown": "@hidehisaarai1213  i was not part of previous competition i did thought of using long er duration but problem is how we break predictions  and assign them to 5 sec clips."
        }
      ]
    },
    {
      "id": 1270346,
      "postDate": "2021-04-11T14:51:26.187Z",
      "content": "<p>Thanks a lot for your efforts !</p>\n<p>I was asking myself, how you choose the spectrogram parameters and how they influence prediction results.<br>\nCould you provide us an intuition on this ? How did for example choose the frequency range? I saw different kernels using different parameters.</p>",
      "rawMarkdown": "Thanks a lot for your efforts !\n\nI was asking myself, how you choose the spectrogram parameters and how they influence prediction results.\nCould you provide us an intuition on this ? How did for example choose the frequency range? I saw different kernels using different parameters.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1270820,
          "postDate": "2021-04-12T04:06:12.393Z",
          "content": "<p>Thanks, comment.</p>\n<p>I have referred to the previous bird call detection competition's code to determine parameters.<br>\n<a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\" target=\"_blank\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a></p>\n<p>I think this is enough to detect bird calls.</p>",
          "rawMarkdown": "Thanks, comment.\n\nI have referred to the previous bird call detection competition's code to determine parameters.\nhttps://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\n\nI think this is enough to detect bird calls.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1264978,
      "postDate": "2021-04-06T14:29:46.547Z",
      "content": "<p>Maybe better to use <code>np.savez_compressed</code> instead of <code>np.save</code>. I think the former uses compression whereas the latter does not compress the data which leads to large dataset size (and you need to separate the dataset into pieces and pieces) </p>",
      "rawMarkdown": "Maybe better to use `np.savez_compressed` instead of `np.save`. I think the former uses compression whereas the latter does not compress the data which leads to large dataset size (and you need to separate the dataset into pieces and pieces) ",
      "votes": 1,
      "replies": [
        {
          "id": 1265469,
          "postDate": "2021-04-06T23:21:26.453Z",
          "content": "<p>Thank you comment!<br>\nI will try it after make this dataset starter code!</p>",
          "rawMarkdown": "Thank you comment!\nI will try it after make this dataset starter code!"
        },
        {
          "id": 1265793,
          "postDate": "2021-04-07T08:09:57.507Z",
          "content": "<p>I have tried.</p>\n<pre><code>def save_npz_v2(primary_label, filenamee):\n    path = f\"{config.INPUT_ROOT}/train_short_audio/{primary_label}/{filename}\"\n    mel_spec = load_mel_spec(path)\n    return (filename, mel_spec)\n</code></pre>\n<pre><code>mel_specs = joblib.Parallel(n_jobs=8)(\n    joblib.delayed(save_npz_v2)(primary_label, filename)\n    for primary_label, filename in df[[\"primary_label\", \"filename\"]].values\n)\nmel_specs = {i: j for i, j in mel_specs}\nnp.savez_compressed(primary_label, **mel_specs)\n</code></pre>\n<p>However, the reduction was few compared to zip compression.<br>\n(acafly 424235171→424232531)</p>",
          "rawMarkdown": "I have tried.\n\n```python\ndef save_npz_v2(primary_label, filenamee):\n    path = f\"{config.INPUT_ROOT}/train_short_audio/{primary_label}/{filename}\"\n    mel_spec = load_mel_spec(path)\n    return (filename, mel_spec)\n```\n\n```python\nmel_specs = joblib.Parallel(n_jobs=8)(\n    joblib.delayed(save_npz_v2)(primary_label, filename)\n    for primary_label, filename in df[[\"primary_label\", \"filename\"]].values\n)\nmel_specs = {i: j for i, j in mel_specs}\nnp.savez_compressed(primary_label, **mel_specs)\n   \n```\n\nHowever, the reduction was few compared to zip compression.\n(acafly 424235171→424232531)",
          "votes": 1
        },
        {
          "id": 1265930,
          "postDate": "2021-04-07T10:48:40.270Z",
          "content": "<p>Thanks for the try! Does that script compressing for each species? I think compressing for each audio clip may have some differences</p>",
          "rawMarkdown": "Thanks for the try! Does that script compressing for each species? I think compressing for each audio clip may have some differences",
          "votes": 2
        },
        {
          "id": 1266065,
          "postDate": "2021-04-07T13:02:32.130Z",
          "content": "<p>Thanks, I will try!</p>\n<p>The reduction is important for me because this competition's dataset is too large to treat.<br>\nI'm trying another reduction idea now, if it will get a good result, I share.</p>",
          "rawMarkdown": "Thanks, I will try!\n\nThe reduction is important for me because this competition's dataset is too large to treat.\nI'm trying another reduction idea now, if it will get a good result, I share."
        },
        {
          "id": 1266071,
          "postDate": "2021-04-07T13:06:14.280Z",
          "content": "<blockquote>\n  <p>this competition's dataset is too large to treat.</p>\n</blockquote>\n<p>Yes indeed. and we also have the Cornell Birdcall's datasets as well.</p>",
          "rawMarkdown": "> this competition's dataset is too large to treat.\n\nYes indeed. and we also have the Cornell Birdcall's datasets as well.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1273697,
      "postDate": "2021-04-14T14:53:48.217Z",
      "content": "<p>My npy data is almost 105G in total. And I think we use the same code to generate the mels. I don't understand why there is such a big difference.</p>",
      "rawMarkdown": "My npy data is almost 105G in total. And I think we use the same code to generate the mels. I don't understand why there is such a big difference.",
      "replies": [
        {
          "id": 1320428,
          "postDate": "2021-05-24T04:48:19.490Z",
          "content": "<p>convert data type to uint8 if  u no doing it</p>",
          "rawMarkdown": "convert data type to uint8 if  u no doing it",
          "votes": 2
        }
      ]
    },
    {
      "id": 1270819,
      "postDate": "2021-04-12T04:05:31.627Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1300494,
      "postDate": "2021-05-10T13:58:48.347Z",
      "content": "<p>Thanks for your sharing.</p>",
      "rawMarkdown": "Thanks for your sharing.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1271555,
      "author_name": "P.Zhao",
      "author_url": "",
      "post_date": "2021-04-12T17:33:45.013000",
      "content": "<p>One more slightly stupid question :</p>\n<p>In the test data, we need to predict for chunks of 5s. However, the spectrogram data here is for the whole audio file as far as I understand it ? Wont this be a problem ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1271798,
          "author_name": "Takamichi Toda",
          "author_url": "",
          "post_date": "2021-04-12T23:18:47.640000",
          "content": "<p>Yes, this spectrogram data here is for the whole audio file.<br>\nSo, you need to crop the spectrogram for 5 seconds in training.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1272859,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2021-04-13T20:26:00.633000",
          "content": "<p>Just to lightly seed another line of thought:<br>\nWhen you process the soundscapes your model makes predictions on a 5s window, but does have access to the full file. There may be nice ways to take advantage of longer time windows when predicting a 5s window. I don't /know/ that there's an approach that will work, but I do think it's a bit under-explored.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1273019,
          "author_name": "Takamichi Toda",
          "author_url": "",
          "post_date": "2021-04-14T02:51:06.780000",
          "content": "<p>Thank you for your advice!<br>\nI may be able to make some predictions from the audio before and after.<br>\nI'll challenge!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1273454,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-04-14T11:16:09.043000",
          "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> To your point, the SED model that was heavily used in previous competition and is also reused in some public notebooks here does look at more than 5 second clips.  Don't you agree?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1274020,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2021-04-14T22:47:22.393000",
          "content": "<p>FYI, prediction on 30s window was better than prediction on 5s window for me in the previous competition.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1276734,
          "author_name": "undisclosed",
          "author_url": "",
          "post_date": "2021-04-17T22:36:22.953000",
          "content": "<p>As a birder, the use of longer time segments makes a lot of intuitive sense. There are a lot of cases in which pattern of repetition matters as much or more than the the sound of a single element or phrase. A single element on its own heard at a distance might be indistinguishable between a Red-eyed Vireo, a Blue-headed Vireo, a Yellow-throated Vireo, a Philadelphia Vireo, or even an American Robin. But if the bird keeps singing, incessantly, particularly if is a summer afternoon, then Red-eyed Vireo is a good bet (assuming you’re in its range, of course, and the habitat is right).  Likewise, a single 5 second chunk of a Northern Mockingbird could be really confusing, but as it keeps singing, the mystery vanishes (at least to a listening birder); mockingbirds often repeat elements a few times before switching to a different pattern for a few repetitions, and on and on. Though the content varies, the pattern is reliable. </p>\n<p>I have to admit I’m still learning my way around the options for implementing this,  but my instinct says it would be worth the effort. Thanks to those who have offered up some leads. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1320433,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-05-24T04:52:27.560000",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>  i was not part of previous competition i did thought of using long er duration but problem is how we break predictions  and assign them to 5 sec clips.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1270346,
      "author_name": "P.Zhao",
      "author_url": "",
      "post_date": "2021-04-11T14:51:26.187000",
      "content": "<p>Thanks a lot for your efforts !</p>\n<p>I was asking myself, how you choose the spectrogram parameters and how they influence prediction results.<br>\nCould you provide us an intuition on this ? How did for example choose the frequency range? I saw different kernels using different parameters.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1270820,
          "author_name": "Takamichi Toda",
          "author_url": "",
          "post_date": "2021-04-12T04:06:12.393000",
          "content": "<p>Thanks, comment.</p>\n<p>I have referred to the previous bird call detection competition's code to determine parameters.<br>\n<a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\" target=\"_blank\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a></p>\n<p>I think this is enough to detect bird calls.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1264978,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2021-04-06T14:29:46.547000",
      "content": "<p>Maybe better to use <code>np.savez_compressed</code> instead of <code>np.save</code>. I think the former uses compression whereas the latter does not compress the data which leads to large dataset size (and you need to separate the dataset into pieces and pieces) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1265469,
          "author_name": "Takamichi Toda",
          "author_url": "",
          "post_date": "2021-04-06T23:21:26.453000",
          "content": "<p>Thank you comment!<br>\nI will try it after make this dataset starter code!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1265793,
          "author_name": "Takamichi Toda",
          "author_url": "",
          "post_date": "2021-04-07T08:09:57.507000",
          "content": "<p>I have tried.</p>\n<pre><code>def save_npz_v2(primary_label, filenamee):\n    path = f\"{config.INPUT_ROOT}/train_short_audio/{primary_label}/{filename}\"\n    mel_spec = load_mel_spec(path)\n    return (filename, mel_spec)\n</code></pre>\n<pre><code>mel_specs = joblib.Parallel(n_jobs=8)(\n    joblib.delayed(save_npz_v2)(primary_label, filename)\n    for primary_label, filename in df[[\"primary_label\", \"filename\"]].values\n)\nmel_specs = {i: j for i, j in mel_specs}\nnp.savez_compressed(primary_label, **mel_specs)\n</code></pre>\n<p>However, the reduction was few compared to zip compression.<br>\n(acafly 424235171→424232531)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1265930,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2021-04-07T10:48:40.270000",
          "content": "<p>Thanks for the try! Does that script compressing for each species? I think compressing for each audio clip may have some differences</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1266065,
          "author_name": "Takamichi Toda",
          "author_url": "",
          "post_date": "2021-04-07T13:02:32.130000",
          "content": "<p>Thanks, I will try!</p>\n<p>The reduction is important for me because this competition's dataset is too large to treat.<br>\nI'm trying another reduction idea now, if it will get a good result, I share.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1266071,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2021-04-07T13:06:14.280000",
          "content": "<blockquote>\n  <p>this competition's dataset is too large to treat.</p>\n</blockquote>\n<p>Yes indeed. and we also have the Cornell Birdcall's datasets as well.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1273697,
      "author_name": "onlytwow",
      "author_url": "",
      "post_date": "2021-04-14T14:53:48.217000",
      "content": "<p>My npy data is almost 105G in total. And I think we use the same code to generate the mels. I don't understand why there is such a big difference.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1320428,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-05-24T04:48:19.490000",
          "content": "<p>convert data type to uint8 if  u no doing it</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1270819,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-12T04:05:31.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1300494,
      "author_name": "majunfu",
      "author_url": "",
      "post_date": "2021-05-10T13:58:48.347000",
      "content": "<p>Thanks for your sharing.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1264163": "I think this competition needs lots of computing time to calculate the spectrograms.\nWhen we pre-compute the spectrograms, we can save computing time.\n  \nI publish pre-computed spectrograms dataset.  \nThe dataset was very large, I split it by the initial alphabet.\n\n- [BirdCLEF Spectrogram 01 (a)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-01-a)\n- [BirdCLEF Spectrogram 02 (b)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-02-b)\n- [BirdCLEF Spectrogram 03 (c)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-03-c)\n- [BirdCLEF Spectrogram 04 (d-g)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-04-dg)\n- [BirdCLEF Spectrogram 05 (h-m)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-05-hm)\n- [BirdCLEF Spectrogram 06 (n-p)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-06-np)\n- [BirdCLEF Spectrogram 07 (q-s)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-07-qs)\n- [BirdCLEF Spectrogram 08 (t-z)](https://www.kaggle.com/takamichitoda/birdclef-spectrogram-08-tz)\n\n  \nI calculate spectrograms by using this code.\n\n```python\nclass config:\n    INPUT_ROOT = \"/kaggle/input/birdclef-2021\"\n    WORK_ROOT = \"/kaggle/working\"\n    FMIN = 20\n    FMAX = 16000\n    SPEC_HEIGHT = 128\n    N_FFT = 2048\n```\n```python\ndef load_mel_spec(path):\n    data, samplerate = sf.read(path)\n    mel_spec = librosa.feature.melspectrogram(y=data, \n                                              sr=samplerate, \n                                              n_fft=config.N_FFT, \n                                              n_mels=config.SPEC_HEIGHT, \n                                              fmin=config.FMIN, \n                                              fmax=config.FMAX)\n    mel_spec = librosa.power_to_db(mel_spec, ref=np.max) \n    return mel_spec\n```\n```python\ndef save_npz(primary_label, filename):\n    path = f\"{config.INPUT_ROOT}/train_short_audio/{primary_label}/{filename}\"\n    mel_spec = load_mel_spec(path)\n    np.save(f\"{config.WORK_ROOT}/{primary_label}/{filename}\", mel_spec)\n```\n\n\nI will make the starter code with this dataset.\n\n",
    "1271555": "One more slightly stupid question :\n\nIn the test data, we need to predict for chunks of 5s. However, the spectrogram data here is for the whole audio file as far as I understand it ? Wont this be a problem ?",
    "1270346": "Thanks a lot for your efforts !\n\nI was asking myself, how you choose the spectrogram parameters and how they influence prediction results.\nCould you provide us an intuition on this ? How did for example choose the frequency range? I saw different kernels using different parameters.\n",
    "1264978": "Maybe better to use `np.savez_compressed` instead of `np.save`. I think the former uses compression whereas the latter does not compress the data which leads to large dataset size (and you need to separate the dataset into pieces and pieces) ",
    "1273697": "My npy data is almost 105G in total. And I think we use the same code to generate the mels. I don't understand why there is such a big difference.",
    "1270819": "",
    "1300494": "Thanks for your sharing."
  }
}