{
  "id": 159001,
  "title": "Dataset converted to spectrograms",
  "url": "/competitions/birdsong-recognition/discussion/159001",
  "author_name": "ryches",
  "post_date": "2020-06-16T04:48:35.384000",
  "votes": 17,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In the past people have success with spectrograms of the audio files. I converted the mp3's with librosa in this kernel <a href=\"https://www.kaggle.com/ryches/birdsong-data-prep-datasetpls?scriptVersionId=36470291\">https://www.kaggle.com/ryches/birdsong-data-prep-datasetpls?scriptVersionId=36470291</a>. </p>\n\n<p>Here is the dataset with the actual images. <a href=\"https://www.kaggle.com/ryches/birdsongspectrograms\">https://www.kaggle.com/ryches/birdsongspectrograms</a></p>\n\n<p>I have a keras kernel training that I will make public once it is finished and I work out how to make predictions on test. </p>",
  "messages": [
    {
      "id": 888042,
      "postDate": "2020-06-16T04:48:35.383Z",
      "content": "<p>In the past people have success with spectrograms of the audio files. I converted the mp3's with librosa in this kernel <a href=\"https://www.kaggle.com/ryches/birdsong-data-prep-datasetpls?scriptVersionId=36470291\">https://www.kaggle.com/ryches/birdsong-data-prep-datasetpls?scriptVersionId=36470291</a>. </p>\n\n<p>Here is the dataset with the actual images. <a href=\"https://www.kaggle.com/ryches/birdsongspectrograms\">https://www.kaggle.com/ryches/birdsongspectrograms</a></p>\n\n<p>I have a keras kernel training that I will make public once it is finished and I work out how to make predictions on test. </p>",
      "rawMarkdown": "In the past people have success with spectrograms of the audio files. I converted the mp3's with librosa in this kernel https://www.kaggle.com/ryches/birdsong-data-prep-datasetpls?scriptVersionId=36470291. \n\nHere is the dataset with the actual images. https://www.kaggle.com/ryches/birdsongspectrograms\n\nI have a keras kernel training that I will make public once it is finished and I work out how to make predictions on test. ",
      "votes": 17
    },
    {
      "id": 888226,
      "postDate": "2020-06-16T07:51:49.080Z",
      "content": "<p>Thanks for taking the time to do it !</p>\n\n<p>In my experiments, loading mp3s is much faster using the <code>audio2numpy</code> package, you should be able to save some time :) </p>\n\n<p>```\n!pip install audio2numpy\nfrom audio2numpy import open_audio</p>\n\n<p>def load_mp3(f):\n    return open_audio(f)[0]\n```</p>",
      "rawMarkdown": "Thanks for taking the time to do it !\n\nIn my experiments, loading mp3s is much faster using the `audio2numpy` package, you should be able to save some time :) \n\n```\n!pip install audio2numpy\nfrom audio2numpy import open_audio\n\ndef load_mp3(f):\n    return open_audio(f)[0]\n```",
      "votes": 3,
      "replies": [
        {
          "id": 888249,
          "postDate": "2020-06-16T08:20:08.600Z",
          "content": "<p>I just tried it out and I am not seeing any measurable gain. I did find that with librosa it is way faster if I define the sampling_rate rather than having it define it itself. Librosa also allows me to clip audio at a certain length upon loading which I dont see as an option in audio2numpy. </p>\n\n<p>Didnt extensively test, but just watching tqdm it seemed about the same. I will let it run overnight and see where it ends up though</p>",
          "rawMarkdown": "I just tried it out and I am not seeing any measurable gain. I did find that with librosa it is way faster if I define the sampling_rate rather than having it define it itself. Librosa also allows me to clip audio at a certain length upon loading which I dont see as an option in audio2numpy. \n\nDidnt extensively test, but just watching tqdm it seemed about the same. I will let it run overnight and see where it ends up though",
          "votes": 1
        },
        {
          "id": 888273,
          "postDate": "2020-06-16T08:50:21.907Z",
          "content": "<p>Indeed, I didn't define the sampling rate for librosa in my benchmarks, thanks for pointing that out !</p>",
          "rawMarkdown": "Indeed, I didn't define the sampling rate for librosa in my benchmarks, thanks for pointing that out !",
          "votes": 1
        }
      ]
    },
    {
      "id": 888634,
      "postDate": "2020-06-16T13:29:22.287Z",
      "content": "<p>I think you need to do resampling on the audio clips to some fixed sampling rate (let's say 22050 Hz for example and discard clips with lower sr) before converting it to mel-spectrogram, otherwise the scale of the yaxis will not much between audio clips with different sampling rate.</p>",
      "rawMarkdown": "I think you need to do resampling on the audio clips to some fixed sampling rate (let's say 22050 Hz for example and discard clips with lower sr) before converting it to mel-spectrogram, otherwise the scale of the yaxis will not much between audio clips with different sampling rate.",
      "votes": 1
    },
    {
      "id": 888981,
      "postDate": "2020-06-16T17:37:00.270Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryches/birdsong-keras-starter?scriptVersionId=36533238\">https://www.kaggle.com/ryches/birdsong-keras-starter?scriptVersionId=36533238</a></p>\n\n<p>starting to do some training. Still trying to figure out exactly how to apply to this test dataset in the dark. I am making some guesses as to how it will be organized. </p>",
      "rawMarkdown": "https://www.kaggle.com/ryches/birdsong-keras-starter?scriptVersionId=36533238\n\nstarting to do some training. Still trying to figure out exactly how to apply to this test dataset in the dark. I am making some guesses as to how it will be organized. ",
      "votes": 2
    },
    {
      "id": 888142,
      "postDate": "2020-06-16T06:50:17.443Z",
      "content": "<p>Thanks, <a href=\"/ryches\">@ryches</a>! \nI'm getting a 404 on the dataset link. I think it might be private?</p>",
      "rawMarkdown": "Thanks, @ryches! \nI'm getting a 404 on the dataset link. I think it might be private?",
      "votes": 2,
      "replies": [
        {
          "id": 888168,
          "postDate": "2020-06-16T07:00:43.283Z",
          "content": "<p>Thanks for mentioning this. Just fixed it. The dataset is a bit borked. I am fixing it now. When I cast it to np.uint8 I lost quite a bit of information. Rerunning now</p>",
          "rawMarkdown": "Thanks for mentioning this. Just fixed it. The dataset is a bit borked. I am fixing it now. When I cast it to np.uint8 I lost quite a bit of information. Rerunning now",
          "votes": 2
        }
      ]
    },
    {
      "id": 888620,
      "postDate": "2020-06-16T13:19:49.583Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 888226,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-06-16T07:51:49.080000",
      "content": "<p>Thanks for taking the time to do it !</p>\n\n<p>In my experiments, loading mp3s is much faster using the <code>audio2numpy</code> package, you should be able to save some time :) </p>\n\n<p>```\n!pip install audio2numpy\nfrom audio2numpy import open_audio</p>\n\n<p>def load_mp3(f):\n    return open_audio(f)[0]\n```</p>",
      "votes": 3,
      "replies": [
        {
          "id": 888249,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-06-16T08:20:08.600000",
          "content": "<p>I just tried it out and I am not seeing any measurable gain. I did find that with librosa it is way faster if I define the sampling_rate rather than having it define it itself. Librosa also allows me to clip audio at a certain length upon loading which I dont see as an option in audio2numpy. </p>\n\n<p>Didnt extensively test, but just watching tqdm it seemed about the same. I will let it run overnight and see where it ends up though</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 888273,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-06-16T08:50:21.907000",
          "content": "<p>Indeed, I didn't define the sampling rate for librosa in my benchmarks, thanks for pointing that out !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 888634,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2020-06-16T13:29:22.287000",
      "content": "<p>I think you need to do resampling on the audio clips to some fixed sampling rate (let's say 22050 Hz for example and discard clips with lower sr) before converting it to mel-spectrogram, otherwise the scale of the yaxis will not much between audio clips with different sampling rate.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 888981,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-06-16T17:37:00.270000",
      "content": "<p><a href=\"https://www.kaggle.com/ryches/birdsong-keras-starter?scriptVersionId=36533238\">https://www.kaggle.com/ryches/birdsong-keras-starter?scriptVersionId=36533238</a></p>\n\n<p>starting to do some training. Still trying to figure out exactly how to apply to this test dataset in the dark. I am making some guesses as to how it will be organized. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 888142,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2020-06-16T06:50:17.443000",
      "content": "<p>Thanks, <a href=\"/ryches\">@ryches</a>! \nI'm getting a 404 on the dataset link. I think it might be private?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 888168,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-06-16T07:00:43.283000",
          "content": "<p>Thanks for mentioning this. Just fixed it. The dataset is a bit borked. I am fixing it now. When I cast it to np.uint8 I lost quite a bit of information. Rerunning now</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 888620,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-16T13:19:49.583000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "888042": "In the past people have success with spectrograms of the audio files. I converted the mp3's with librosa in this kernel https://www.kaggle.com/ryches/birdsong-data-prep-datasetpls?scriptVersionId=36470291. \n\nHere is the dataset with the actual images. https://www.kaggle.com/ryches/birdsongspectrograms\n\nI have a keras kernel training that I will make public once it is finished and I work out how to make predictions on test. ",
    "888226": "Thanks for taking the time to do it !\n\nIn my experiments, loading mp3s is much faster using the `audio2numpy` package, you should be able to save some time :) \n\n```\n!pip install audio2numpy\nfrom audio2numpy import open_audio\n\ndef load_mp3(f):\n    return open_audio(f)[0]\n```",
    "888634": "I think you need to do resampling on the audio clips to some fixed sampling rate (let's say 22050 Hz for example and discard clips with lower sr) before converting it to mel-spectrogram, otherwise the scale of the yaxis will not much between audio clips with different sampling rate.",
    "888981": "https://www.kaggle.com/ryches/birdsong-keras-starter?scriptVersionId=36533238\n\nstarting to do some training. Still trying to figure out exactly how to apply to this test dataset in the dark. I am making some guesses as to how it will be organized. ",
    "888142": "Thanks, @ryches! \nI'm getting a 404 on the dataset link. I think it might be private?",
    "888620": ""
  }
}