{
  "id": 175616,
  "title": "A script for you to resample your data.",
  "url": "/competitions/birdsong-recognition/discussion/175616",
  "author_name": "",
  "post_date": "2020-08-18T19:24:09.800249700Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>I'd like to provide the code I wrote that will resample and save all of your audio files as <code>.wav</code> at a sampling rate of your choice. It's set up to use multiprocessing and will use all but one available CPU. Expect 8-10 hours on 8 CPUs. </p>\n<p>You will need to create your own <code>if __name__ == '__main__'</code> entry point and call: </p>\n<p><code>resample_all(old_loc, new_loc, sr, restart, ext, manifest)</code></p>\n<p><code>old_loc</code>: the 'train_audio' folder, and <code>new_loc</code> is any other location (can be the same as the old). </p>\n<p><code>sr</code>: The desired sampling rate for resampling.</p>\n<p><code>restart (default False)</code>: If True, the script will look for a manifest containing all files have already been resampled. </p>\n<p><code>ext ('WAV' default)</code>: Don't mess with this. I didn't set it up to accept any other file types. </p>\n<p><code>manifest (default None)</code>: Tells the script where to save a log of the files it's completed, so it doesn't have to start over from scratch if it is interrupted. Mandatory argument if <code>restart=True</code>  <strong><em>Corrupted data exists in this dataset and will interrupt the process, so using this is a good idea</em></strong></p>\n<p>Last thing: The method expects that <code>old_loc</code> be a directory with the traditional labeling format of image data:</p>\n<pre><code>training_audio/ # old location\n        ├─── Label1/\n        │       └─file1.m4a\n        └───Label2/\n                 └─file2.mp3\n</code></pre>\n<p>Okay, here ya go. </p>\n<pre><code># File and I/O\nimport warnings\nimport os\n\n# Audio I/O and processing\nimport librosa\nimport soundfile as sf\n\n# Multi-processing tools\nfrom multiprocessing import Pool\n\n\ndef resample(old_path, new_path, sr, ext='WAV', manifest=None):\n\n    with warnings.catch_warnings():\n        warnings.simplefilter(\"ignore\", UserWarning)\n        audio, sr = librosa.load(old_path, sr=sr)\n\n    os.makedirs(os.path.dirname(new_path), exist_ok=True)\n    with sf.SoundFile(new_path, 'w', sr, channels=1, format=ext) as f:\n        f.write(audio)\n\n    # write filename to manifest as completed.\n    fname = os.path.basename(old_path)\n    if manifest is not None:\n        with open(manifest, 'a') as manifest:\n            manifest.write(fname + '\\n')\n\n\ndef rename(file_path, ext='.wav'):\n    return 'rs_' + os.path.splitext(file_path)[0] + ext\n\n\ndef resample_all(old_loc, new_loc, sr, restart=False, manifest=None):\n\n    if restart:\n        with open(manifest) as m:\n            completed = set([line.strip() for line in m])\n\n        folders = [d for d in os.scandir(old_loc) if os.path.isdir(d.path)]\n\n    for i, folder in enumerate(folders):\n        print(f'Working on folder {folder.name}. Overall progress: {int(100 * i / len(folders))}%   \\r', end=\"\")\n        dirname = folder.name\n\n        if restart:\n            files = [f for f in os.scandir(folder) if (f.name not in completed) and os.path.isfile(f.path)]\n        else:\n            files = [f for f in os.scandir(folder) if os.path.isfile(f.path)]\n\n        fpaths = [f.path for f in files]\n        fnames = [f.name for f in files]\n        renamed = map(rename, fnames)\n\n        new_paths = [os.path.join(new_loc, dirname, f) for f in renamed]\n\n        with Pool(os.cpu_count() - 1) as p:\n            N = len(fpaths)\n            p.starmap(resample, zip(fpaths, new_paths, [sr] * N, ['WAV'] * N, [manifest] * N))\n</code></pre>",
  "messages": [
    {
      "id": "976309",
      "postDate": "08/18/2020 19:24:09",
      "content": "<p>Hi all!</p>\n<p>I'd like to provide the code I wrote that will resample and save all of your audio files as <code>.wav</code> at a sampling rate of your choice. It's set up to use multiprocessing and will use all but one available CPU. Expect 8-10 hours on 8 CPUs. </p>\n<p>You will need to create your own <code>if __name__ == '__main__'</code> entry point and call: </p>\n<p><code>resample_all(old_loc, new_loc, sr, restart, ext, manifest)</code></p>\n<p><code>old_loc</code>: the 'train_audio' folder, and <code>new_loc</code> is any other location (can be the same as the old). </p>\n<p><code>sr</code>: The desired sampling rate for resampling.</p>\n<p><code>restart (default False)</code>: If True, the script will look for a manifest containing all files have already been resampled. </p>\n<p><code>ext ('WAV' default)</code>: Don't mess with this. I didn't set it up to accept any other file types. </p>\n<p><code>manifest (default None)</code>: Tells the script where to save a log of the files it's completed, so it doesn't have to start over from scratch if it is interrupted. Mandatory argument if <code>restart=True</code>  <strong><em>Corrupted data exists in this dataset and will interrupt the process, so using this is a good idea</em></strong></p>\n<p>Last thing: The method expects that <code>old_loc</code> be a directory with the traditional labeling format of image data:</p>\n<pre><code>training_audio/ # old location\n        ├─── Label1/\n        │       └─file1.m4a\n        └───Label2/\n                 └─file2.mp3\n</code></pre>\n<p>Okay, here ya go. </p>\n<pre><code># File and I/O\nimport warnings\nimport os\n\n# Audio I/O and processing\nimport librosa\nimport soundfile as sf\n\n# Multi-processing tools\nfrom multiprocessing import Pool\n\n\ndef resample(old_path, new_path, sr, ext='WAV', manifest=None):\n\n    with warnings.catch_warnings():\n        warnings.simplefilter(\"ignore\", UserWarning)\n        audio, sr = librosa.load(old_path, sr=sr)\n\n    os.makedirs(os.path.dirname(new_path), exist_ok=True)\n    with sf.SoundFile(new_path, 'w', sr, channels=1, format=ext) as f:\n        f.write(audio)\n\n    # write filename to manifest as completed.\n    fname = os.path.basename(old_path)\n    if manifest is not None:\n        with open(manifest, 'a') as manifest:\n            manifest.write(fname + '\\n')\n\n\ndef rename(file_path, ext='.wav'):\n    return 'rs_' + os.path.splitext(file_path)[0] + ext\n\n\ndef resample_all(old_loc, new_loc, sr, restart=False, manifest=None):\n\n    if restart:\n        with open(manifest) as m:\n            completed = set([line.strip() for line in m])\n\n        folders = [d for d in os.scandir(old_loc) if os.path.isdir(d.path)]\n\n    for i, folder in enumerate(folders):\n        print(f'Working on folder {folder.name}. Overall progress: {int(100 * i / len(folders))}%   \\r', end=\"\")\n        dirname = folder.name\n\n        if restart:\n            files = [f for f in os.scandir(folder) if (f.name not in completed) and os.path.isfile(f.path)]\n        else:\n            files = [f for f in os.scandir(folder) if os.path.isfile(f.path)]\n\n        fpaths = [f.path for f in files]\n        fnames = [f.name for f in files]\n        renamed = map(rename, fnames)\n\n        new_paths = [os.path.join(new_loc, dirname, f) for f in renamed]\n\n        with Pool(os.cpu_count() - 1) as p:\n            N = len(fpaths)\n            p.starmap(resample, zip(fpaths, new_paths, [sr] * N, ['WAV'] * N, [manifest] * N))\n</code></pre>",
      "rawMarkdown": "Hi all!\n\nI'd like to provide the code I wrote that will resample and save all of your audio files as `.wav` at a sampling rate of your choice. It's set up to use multiprocessing and will use all but one available CPU. Expect 8-10 hours on 8 CPUs. \n\nYou will need to create your own `if __name__ == '__main__'` entry point and call: \n\n`resample_all(old_loc, new_loc, sr, restart, ext, manifest)`\n\n`old_loc`: the 'train_audio' folder, and `new_loc` is any other location (can be the same as the old). \n\n`sr`: The desired sampling rate for resampling.\n\n`restart (default False)`: If True, the script will look for a manifest containing all files have already been resampled. \n\n`ext ('WAV' default)`: Don't mess with this. I didn't set it up to accept any other file types. \n\n`manifest (default None)`: Tells the script where to save a log of the files it's completed, so it doesn't have to start over from scratch if it is interrupted. Mandatory argument if `restart=True`  ***Corrupted data exists in this dataset and will interrupt the process, so using this is a good idea***\n\nLast thing: The method expects that `old_loc` be a directory with the traditional labeling format of image data:\n\n```\ntraining_audio/ # old location\n        ├─── Label1/\n        │       └─file1.m4a\n        └───Label2/\n                 └─file2.mp3\n```\n\nOkay, here ya go. \n```\n# File and I/O\nimport warnings\nimport os\n\n# Audio I/O and processing\nimport librosa\nimport soundfile as sf\n\n# Multi-processing tools\nfrom multiprocessing import Pool\n\n\ndef resample(old_path, new_path, sr, ext='WAV', manifest=None):\n\n    with warnings.catch_warnings():\n        warnings.simplefilter(\"ignore\", UserWarning)\n        audio, sr = librosa.load(old_path, sr=sr)\n\n    os.makedirs(os.path.dirname(new_path), exist_ok=True)\n    with sf.SoundFile(new_path, 'w', sr, channels=1, format=ext) as f:\n        f.write(audio)\n\n    # write filename to manifest as completed.\n    fname = os.path.basename(old_path)\n    if manifest is not None:\n        with open(manifest, 'a') as manifest:\n            manifest.write(fname + '\\n')\n\n\ndef rename(file_path, ext='.wav'):\n    return 'rs_' + os.path.splitext(file_path)[0] + ext\n\n\ndef resample_all(old_loc, new_loc, sr, restart=False, manifest=None):\n \n    if restart:\n        with open(manifest) as m:\n            completed = set([line.strip() for line in m])\n\n        folders = [d for d in os.scandir(old_loc) if os.path.isdir(d.path)]\n\n    for i, folder in enumerate(folders):\n        print(f'Working on folder {folder.name}. Overall progress: {int(100 * i / len(folders))}%   \\r', end=\"\")\n        dirname = folder.name\n\n        if restart:\n            files = [f for f in os.scandir(folder) if (f.name not in completed) and os.path.isfile(f.path)]\n        else:\n            files = [f for f in os.scandir(folder) if os.path.isfile(f.path)]\n\n        fpaths = [f.path for f in files]\n        fnames = [f.name for f in files]\n        renamed = map(rename, fnames)\n\n        new_paths = [os.path.join(new_loc, dirname, f) for f in renamed]\n\n        with Pool(os.cpu_count() - 1) as p:\n            N = len(fpaths)\n            p.starmap(resample, zip(fpaths, new_paths, [sr] * N, ['WAV'] * N, [manifest] * N))\n\n```",
      "votes": null
    },
    {
      "id": "977762",
      "postDate": "08/19/2020 17:34:40",
      "content": "<p>Thanks! There is a corresponding information that has been posted already:</p>\n<ul>\n<li>resamples data: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/164197\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/164197</a></li>\n<li>resampling code: <a href=\"https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast\" target=\"_blank\">https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast</a> (first commented cell)</li>\n</ul>",
      "rawMarkdown": "Thanks! There is a corresponding information that has been posted already:\n- resamples data: https://www.kaggle.com/c/birdsong-recognition/discussion/164197\n- resampling code: https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast (first commented cell)",
      "votes": null
    },
    {
      "id": "978065",
      "postDate": "08/19/2020 22:48:09",
      "content": "<p>what is benefit of saving it is as wav file? <br>\nwouldn't you get the same data as librosa.load(old_path, sr=sr)?</p>\n<p>I saw a lot of notebooks first saving as wav then saving as melspectrogram as colored png.<br>\nYou can do grayscale png and have all the data in under 1 GB .  <br>\nAll the key information is preserved in those pngs, i.e. I am not able to tell the difference between reconstructed audio without listening for it very closely or looking at data directy</p>",
      "rawMarkdown": "what is benefit of saving it is as wav file? \nwouldn't you get the same data as librosa.load(old_path, sr=sr)?\n\nI saw a lot of notebooks first saving as wav then saving as melspectrogram as colored png.\nYou can do grayscale png and have all the data in under 1 GB .  \nAll the key information is preserved in those pngs, i.e. I am not able to tell the difference between reconstructed audio without listening for it very closely or looking at data directy",
      "votes": null
    },
    {
      "id": "978080",
      "postDate": "08/19/2020 23:25:34",
      "content": "<p>Because the point of the script is to resample ALL the files in the <code>train_audio</code> directory. It's easier to save them all as one thing, and <code>SoundFile.write()</code> does not have a format argument for mp3. you can check this by calling <code>soundfile.available_formats()</code></p>\n<p>As to your second question, I don't know what you are asking. If it is why I have my own <code>resample</code> method, then please look at that method and see that it indeed resamples by using <code>librosa.load()</code>, just as you mention. That's not the point. The point is that my method also saves resampled files to disk. </p>\n<p>I'm not sure about your melspectrograms comment, as it seems entirely unrelated to this post or my script. </p>\n<p>Please feel free not to use this. I see you are 15th in the competition, so I am very confused by all of your questions. </p>",
      "rawMarkdown": "Because the point of the script is to resample ALL the files in the `train_audio` directory. It's easier to save them all as one thing, and `SoundFile.write()` does not have a format argument for mp3. you can check this by calling `soundfile.available_formats()`\n\nAs to your second question, I don't know what you are asking. If it is why I have my own `resample` method, then please look at that method and see that it indeed resamples by using `librosa.load()`, just as you mention. That's not the point. The point is that my method also saves resampled files to disk. \n\nI'm not sure about your melspectrograms comment, as it seems entirely unrelated to this post or my script. \n\nPlease feel free not to use this. I see you are 15th in the competition, so I am very confused by all of your questions.",
      "votes": null
    },
    {
      "id": "978083",
      "postDate": "08/19/2020 23:29:57",
      "content": "<p>As you can see, I did not post any links to any resampled dataset. The point of my post was to provide users code to resample the data as they see fit. </p>\n<p>Second,  I don't know why you posted a link to a baseline model. This post has nothing to do with that. Again, this post simply gives code for users who want to resample their data to a rate of their choice. </p>\n<p>If the data you want is already available, please feel free to use it. </p>",
      "rawMarkdown": "As you can see, I did not post any links to any resampled dataset. The point of my post was to provide users code to resample the data as they see fit. \n\nSecond,  I don't know why you posted a link to a baseline model. This post has nothing to do with that. Again, this post simply gives code for users who want to resample their data to a rate of their choice. \n\nIf the data you want is already available, please feel free to use it.",
      "votes": null
    },
    {
      "id": "978087",
      "postDate": "08/19/2020 23:34:25",
      "content": "<p>Hello again!</p>\n<p>I'm getting some strange comments on my post, that seem to have nothing to do with the discussion at all.</p>\n<p>I realize there are a few datasets out there that have already been resampled. By all means, use them. That is NOT the point of this discussion. If you already have what you need, move on. No need to leave unrelated remarks and critiques. </p>\n<p>If you want to see one of many ways to resample your data, then you can check out my code. If not, great. I enjoyed writing the code, so I thought I'd share it. That's it. </p>\n<p>Cheers, and good luck!<br>\nWes</p>",
      "rawMarkdown": "Hello again!\n\nI'm getting some strange comments on my post, that seem to have nothing to do with the discussion at all.\n\nI realize there are a few datasets out there that have already been resampled. By all means, use them. That is NOT the point of this discussion. If you already have what you need, move on. No need to leave unrelated remarks and critiques. \n\nIf you want to see one of many ways to resample your data, then you can check out my code. If not, great. I enjoyed writing the code, so I thought I'd share it. That's it. \n\nCheers, and good luck!\nWes",
      "votes": null
    },
    {
      "id": "979213",
      "postDate": "08/20/2020 17:39:47",
      "content": "<p>Thank you for reply.  <br>\n It is my first kaggle, first time working with audio data so I am a bit puzzled at some of the practices here.  </p>\n<p>In no way this is a cretique of the code or the method, simply a question of why would one want to resample? I am asking you since you did it and I saw your post, I also saw it in many other notebooks so it made me even more curious.</p>\n<p>If wav are coming from different sampling rates you are not going to get any different data than one from librosa.load(old_path, sr=sr) so why decode?  I am mainly asking because maybe I am missing something.  At the same time I am guess I am suggesting shorter workflow: get the data you use from mp3s and create a cache for it.</p>\n<p>To me it seems mysterious why people most workflow includes first taking smaller by size mp3 files encoding them in larger wav files, then when running a model only using much smaller amount of data.</p>",
      "rawMarkdown": "Thank you for reply.  \n It is my first kaggle, first time working with audio data so I am a bit puzzled at some of the practices here.  \n\nIn no way this is a cretique of the code or the method, simply a question of why would one want to resample? I am asking you since you did it and I saw your post, I also saw it in many other notebooks so it made me even more curious.\n\nIf wav are coming from different sampling rates you are not going to get any different data than one from librosa.load(old_path, sr=sr) so why decode?  I am mainly asking because maybe I am missing something.  At the same time I am guess I am suggesting shorter workflow: get the data you use from mp3s and create a cache for it.\n\nTo me it seems mysterious why people most workflow includes first taking smaller by size mp3 files encoding them in larger wav files, then when running a model only using much smaller amount of data.",
      "votes": null
    },
    {
      "id": "979277",
      "postDate": "08/20/2020 18:24:27",
      "content": "<p>Well, I'd love to leave them as mp3 files for sure! I just couldn't figure out how to make the <code>soundfile</code> library write <code>mp3</code> format. </p>\n<p>This is also my first real Kaggle and is also my first real ML project.  I am very much out of my depth here. </p>\n<p>Now I'm trying to figure out how to clip the audio files to be all the same length… But some files are silent for the first whole minute!</p>",
      "rawMarkdown": "Well, I'd love to leave them as mp3 files for sure! I just couldn't figure out how to make the `soundfile` library write `mp3` format. \n\nThis is also my first real Kaggle and is also my first real ML project.  I am very much out of my depth here. \n\nNow I'm trying to figure out how to clip the audio files to be all the same length... But some files are silent for the first whole minute!",
      "votes": null
    },
    {
      "id": "979303",
      "postDate": "08/20/2020 18:47:35",
      "content": "<p>in terms of audio, I am in about same situation.  <br>\nIs there any reason soundfile advantageous to librosa? </p>\n<p>My key question was not why resample as wav but why resample at all if you can always resample with librosa when opening? <br>\nI guess I can see one used case if you use some previously written scripts that take in whole directories with same file structure and learn from it.  Tensorflow and torch both have data readers like that but you can always modify them a little bit and be able to do preprocessing such as librosa.load(file,sr=sr).</p>",
      "rawMarkdown": "in terms of audio, I am in about same situation.  \nIs there any reason soundfile advantageous to librosa? \n\nMy key question was not why resample as wav but why resample at all if you can always resample with librosa when opening? \nI guess I can see one used case if you use some previously written scripts that take in whole directories with same file structure and learn from it.  Tensorflow and torch both have data readers like that but you can always modify them a little bit and be able to do preprocessing such as librosa.load(file,sr=sr).",
      "votes": null
    },
    {
      "id": "979355",
      "postDate": "08/20/2020 19:21:05",
      "content": "<p>Well, I suppose I thought that it would be easier to resample all the data at once. I think I might be a lemming. I really don't have a good reason for that except that I saw that other people had done the same. </p>\n<p>Now, I'm even more confused because I see that the test data is 10 minutes long and every 5 second increment has to be labeled. What's more, it looks like it's a multi-classification problem, which I've never done. </p>\n<p>I was going to pre-process all the data to a single length by either padding or clipping every file. That's not an option now that I know how the test data works. </p>\n<p>I don't think I'm going to figure all of this ou til well past the deadline. This what I get for joining a non-beginner's competition!</p>",
      "rawMarkdown": "Well, I suppose I thought that it would be easier to resample all the data at once. I think I might be a lemming. I really don't have a good reason for that except that I saw that other people had done the same. \n\nNow, I'm even more confused because I see that the test data is 10 minutes long and every 5 second increment has to be labeled. What's more, it looks like it's a multi-classification problem, which I've never done. \n\nI was going to pre-process all the data to a single length by either padding or clipping every file. That's not an option now that I know how the test data works. \n\nI don't think I'm going to figure all of this ou til well past the deadline. This what I get for joining a non-beginner's competition!",
      "votes": null
    },
    {
      "id": "979485",
      "postDate": "08/20/2020 22:02:24",
      "content": "<p><a href=\"https://www.kaggle.com/leodav\" target=\"_blank\">@leodav</a> I found that <code>soundfile.read</code> was significantly faster than <code>librosa.load</code> when trying to sample everything at 32kHz. That is why I am sticking with resampling.</p>",
      "rawMarkdown": "leodav I found that `soundfile.read` was significantly faster than `librosa.load` when trying to sample everything at 32kHz. That is why I am sticking with resampling.",
      "votes": null
    },
    {
      "id": "979529",
      "postDate": "08/20/2020 23:23:22",
      "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> Thank you that is the first answer on this that makes.  </p>",
      "rawMarkdown": "returnofsputnik Thank you that is the first answer on this that makes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 977762,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "08/19/2020 17:34:40",
      "content": "<p>Thanks! There is a corresponding information that has been posted already:</p>\n<ul>\n<li>resamples data: <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/164197\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/164197</a></li>\n<li>resampling code: <a href=\"https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast\" target=\"_blank\">https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast</a> (first commented cell)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 978083,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/19/2020 23:29:57",
          "content": "<p>As you can see, I did not post any links to any resampled dataset. The point of my post was to provide users code to resample the data as they see fit. </p>\n<p>Second,  I don't know why you posted a link to a baseline model. This post has nothing to do with that. Again, this post simply gives code for users who want to resample their data to a rate of their choice. </p>\n<p>If the data you want is already available, please feel free to use it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 978065,
      "author_name": "leodav",
      "author_url": "",
      "post_date": "08/19/2020 22:48:09",
      "content": "<p>what is benefit of saving it is as wav file? <br>\nwouldn't you get the same data as librosa.load(old_path, sr=sr)?</p>\n<p>I saw a lot of notebooks first saving as wav then saving as melspectrogram as colored png.<br>\nYou can do grayscale png and have all the data in under 1 GB .  <br>\nAll the key information is preserved in those pngs, i.e. I am not able to tell the difference between reconstructed audio without listening for it very closely or looking at data directy</p>",
      "votes": null,
      "replies": [
        {
          "id": 978080,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/19/2020 23:25:34",
          "content": "<p>Because the point of the script is to resample ALL the files in the <code>train_audio</code> directory. It's easier to save them all as one thing, and <code>SoundFile.write()</code> does not have a format argument for mp3. you can check this by calling <code>soundfile.available_formats()</code></p>\n<p>As to your second question, I don't know what you are asking. If it is why I have my own <code>resample</code> method, then please look at that method and see that it indeed resamples by using <code>librosa.load()</code>, just as you mention. That's not the point. The point is that my method also saves resampled files to disk. </p>\n<p>I'm not sure about your melspectrograms comment, as it seems entirely unrelated to this post or my script. </p>\n<p>Please feel free not to use this. I see you are 15th in the competition, so I am very confused by all of your questions. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979213,
          "author_name": "leodav",
          "author_url": "",
          "post_date": "08/20/2020 17:39:47",
          "content": "<p>Thank you for reply.  <br>\n It is my first kaggle, first time working with audio data so I am a bit puzzled at some of the practices here.  </p>\n<p>In no way this is a cretique of the code or the method, simply a question of why would one want to resample? I am asking you since you did it and I saw your post, I also saw it in many other notebooks so it made me even more curious.</p>\n<p>If wav are coming from different sampling rates you are not going to get any different data than one from librosa.load(old_path, sr=sr) so why decode?  I am mainly asking because maybe I am missing something.  At the same time I am guess I am suggesting shorter workflow: get the data you use from mp3s and create a cache for it.</p>\n<p>To me it seems mysterious why people most workflow includes first taking smaller by size mp3 files encoding them in larger wav files, then when running a model only using much smaller amount of data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979277,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/20/2020 18:24:27",
          "content": "<p>Well, I'd love to leave them as mp3 files for sure! I just couldn't figure out how to make the <code>soundfile</code> library write <code>mp3</code> format. </p>\n<p>This is also my first real Kaggle and is also my first real ML project.  I am very much out of my depth here. </p>\n<p>Now I'm trying to figure out how to clip the audio files to be all the same length… But some files are silent for the first whole minute!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979303,
          "author_name": "leodav",
          "author_url": "",
          "post_date": "08/20/2020 18:47:35",
          "content": "<p>in terms of audio, I am in about same situation.  <br>\nIs there any reason soundfile advantageous to librosa? </p>\n<p>My key question was not why resample as wav but why resample at all if you can always resample with librosa when opening? <br>\nI guess I can see one used case if you use some previously written scripts that take in whole directories with same file structure and learn from it.  Tensorflow and torch both have data readers like that but you can always modify them a little bit and be able to do preprocessing such as librosa.load(file,sr=sr).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979355,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/20/2020 19:21:05",
          "content": "<p>Well, I suppose I thought that it would be easier to resample all the data at once. I think I might be a lemming. I really don't have a good reason for that except that I saw that other people had done the same. </p>\n<p>Now, I'm even more confused because I see that the test data is 10 minutes long and every 5 second increment has to be labeled. What's more, it looks like it's a multi-classification problem, which I've never done. </p>\n<p>I was going to pre-process all the data to a single length by either padding or clipping every file. That's not an option now that I know how the test data works. </p>\n<p>I don't think I'm going to figure all of this ou til well past the deadline. This what I get for joining a non-beginner's competition!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979485,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "08/20/2020 22:02:24",
          "content": "<p><a href=\"https://www.kaggle.com/leodav\" target=\"_blank\">@leodav</a> I found that <code>soundfile.read</code> was significantly faster than <code>librosa.load</code> when trying to sample everything at 32kHz. That is why I am sticking with resampling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979529,
          "author_name": "leodav",
          "author_url": "",
          "post_date": "08/20/2020 23:23:22",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> Thank you that is the first answer on this that makes.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 978087,
      "author_name": "wesleyneill",
      "author_url": "",
      "post_date": "08/19/2020 23:34:25",
      "content": "<p>Hello again!</p>\n<p>I'm getting some strange comments on my post, that seem to have nothing to do with the discussion at all.</p>\n<p>I realize there are a few datasets out there that have already been resampled. By all means, use them. That is NOT the point of this discussion. If you already have what you need, move on. No need to leave unrelated remarks and critiques. </p>\n<p>If you want to see one of many ways to resample your data, then you can check out my code. If not, great. I enjoyed writing the code, so I thought I'd share it. That's it. </p>\n<p>Cheers, and good luck!<br>\nWes</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "976309": "Hi all!\n\nI'd like to provide the code I wrote that will resample and save all of your audio files as `.wav` at a sampling rate of your choice. It's set up to use multiprocessing and will use all but one available CPU. Expect 8-10 hours on 8 CPUs. \n\nYou will need to create your own `if __name__ == '__main__'` entry point and call: \n\n`resample_all(old_loc, new_loc, sr, restart, ext, manifest)`\n\n`old_loc`: the 'train_audio' folder, and `new_loc` is any other location (can be the same as the old). \n\n`sr`: The desired sampling rate for resampling.\n\n`restart (default False)`: If True, the script will look for a manifest containing all files have already been resampled. \n\n`ext ('WAV' default)`: Don't mess with this. I didn't set it up to accept any other file types. \n\n`manifest (default None)`: Tells the script where to save a log of the files it's completed, so it doesn't have to start over from scratch if it is interrupted. Mandatory argument if `restart=True`  ***Corrupted data exists in this dataset and will interrupt the process, so using this is a good idea***\n\nLast thing: The method expects that `old_loc` be a directory with the traditional labeling format of image data:\n\n```\ntraining_audio/ # old location\n        ├─── Label1/\n        │       └─file1.m4a\n        └───Label2/\n                 └─file2.mp3\n```\n\nOkay, here ya go. \n```\n# File and I/O\nimport warnings\nimport os\n\n# Audio I/O and processing\nimport librosa\nimport soundfile as sf\n\n# Multi-processing tools\nfrom multiprocessing import Pool\n\n\ndef resample(old_path, new_path, sr, ext='WAV', manifest=None):\n\n    with warnings.catch_warnings():\n        warnings.simplefilter(\"ignore\", UserWarning)\n        audio, sr = librosa.load(old_path, sr=sr)\n\n    os.makedirs(os.path.dirname(new_path), exist_ok=True)\n    with sf.SoundFile(new_path, 'w', sr, channels=1, format=ext) as f:\n        f.write(audio)\n\n    # write filename to manifest as completed.\n    fname = os.path.basename(old_path)\n    if manifest is not None:\n        with open(manifest, 'a') as manifest:\n            manifest.write(fname + '\\n')\n\n\ndef rename(file_path, ext='.wav'):\n    return 'rs_' + os.path.splitext(file_path)[0] + ext\n\n\ndef resample_all(old_loc, new_loc, sr, restart=False, manifest=None):\n \n    if restart:\n        with open(manifest) as m:\n            completed = set([line.strip() for line in m])\n\n        folders = [d for d in os.scandir(old_loc) if os.path.isdir(d.path)]\n\n    for i, folder in enumerate(folders):\n        print(f'Working on folder {folder.name}. Overall progress: {int(100 * i / len(folders))}%   \\r', end=\"\")\n        dirname = folder.name\n\n        if restart:\n            files = [f for f in os.scandir(folder) if (f.name not in completed) and os.path.isfile(f.path)]\n        else:\n            files = [f for f in os.scandir(folder) if os.path.isfile(f.path)]\n\n        fpaths = [f.path for f in files]\n        fnames = [f.name for f in files]\n        renamed = map(rename, fnames)\n\n        new_paths = [os.path.join(new_loc, dirname, f) for f in renamed]\n\n        with Pool(os.cpu_count() - 1) as p:\n            N = len(fpaths)\n            p.starmap(resample, zip(fpaths, new_paths, [sr] * N, ['WAV'] * N, [manifest] * N))\n\n```",
    "977762": "Thanks! There is a corresponding information that has been posted already:\n- resamples data: https://www.kaggle.com/c/birdsong-recognition/discussion/164197\n- resampling code: https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast (first commented cell)",
    "978065": "what is benefit of saving it is as wav file? \nwouldn't you get the same data as librosa.load(old_path, sr=sr)?\n\nI saw a lot of notebooks first saving as wav then saving as melspectrogram as colored png.\nYou can do grayscale png and have all the data in under 1 GB .  \nAll the key information is preserved in those pngs, i.e. I am not able to tell the difference between reconstructed audio without listening for it very closely or looking at data directy",
    "978080": "Because the point of the script is to resample ALL the files in the `train_audio` directory. It's easier to save them all as one thing, and `SoundFile.write()` does not have a format argument for mp3. you can check this by calling `soundfile.available_formats()`\n\nAs to your second question, I don't know what you are asking. If it is why I have my own `resample` method, then please look at that method and see that it indeed resamples by using `librosa.load()`, just as you mention. That's not the point. The point is that my method also saves resampled files to disk. \n\nI'm not sure about your melspectrograms comment, as it seems entirely unrelated to this post or my script. \n\nPlease feel free not to use this. I see you are 15th in the competition, so I am very confused by all of your questions.",
    "978083": "As you can see, I did not post any links to any resampled dataset. The point of my post was to provide users code to resample the data as they see fit. \n\nSecond,  I don't know why you posted a link to a baseline model. This post has nothing to do with that. Again, this post simply gives code for users who want to resample their data to a rate of their choice. \n\nIf the data you want is already available, please feel free to use it.",
    "978087": "Hello again!\n\nI'm getting some strange comments on my post, that seem to have nothing to do with the discussion at all.\n\nI realize there are a few datasets out there that have already been resampled. By all means, use them. That is NOT the point of this discussion. If you already have what you need, move on. No need to leave unrelated remarks and critiques. \n\nIf you want to see one of many ways to resample your data, then you can check out my code. If not, great. I enjoyed writing the code, so I thought I'd share it. That's it. \n\nCheers, and good luck!\nWes",
    "979213": "Thank you for reply.  \n It is my first kaggle, first time working with audio data so I am a bit puzzled at some of the practices here.  \n\nIn no way this is a cretique of the code or the method, simply a question of why would one want to resample? I am asking you since you did it and I saw your post, I also saw it in many other notebooks so it made me even more curious.\n\nIf wav are coming from different sampling rates you are not going to get any different data than one from librosa.load(old_path, sr=sr) so why decode?  I am mainly asking because maybe I am missing something.  At the same time I am guess I am suggesting shorter workflow: get the data you use from mp3s and create a cache for it.\n\nTo me it seems mysterious why people most workflow includes first taking smaller by size mp3 files encoding them in larger wav files, then when running a model only using much smaller amount of data.",
    "979277": "Well, I'd love to leave them as mp3 files for sure! I just couldn't figure out how to make the `soundfile` library write `mp3` format. \n\nThis is also my first real Kaggle and is also my first real ML project.  I am very much out of my depth here. \n\nNow I'm trying to figure out how to clip the audio files to be all the same length... But some files are silent for the first whole minute!",
    "979303": "in terms of audio, I am in about same situation.  \nIs there any reason soundfile advantageous to librosa? \n\nMy key question was not why resample as wav but why resample at all if you can always resample with librosa when opening? \nI guess I can see one used case if you use some previously written scripts that take in whole directories with same file structure and learn from it.  Tensorflow and torch both have data readers like that but you can always modify them a little bit and be able to do preprocessing such as librosa.load(file,sr=sr).",
    "979355": "Well, I suppose I thought that it would be easier to resample all the data at once. I think I might be a lemming. I really don't have a good reason for that except that I saw that other people had done the same. \n\nNow, I'm even more confused because I see that the test data is 10 minutes long and every 5 second increment has to be labeled. What's more, it looks like it's a multi-classification problem, which I've never done. \n\nI was going to pre-process all the data to a single length by either padding or clipping every file. That's not an option now that I know how the test data works. \n\nI don't think I'm going to figure all of this ou til well past the deadline. This what I get for joining a non-beginner's competition!",
    "979485": "leodav I found that `soundfile.read` was significantly faster than `librosa.load` when trying to sample everything at 32kHz. That is why I am sticking with resampling.",
    "979529": "returnofsputnik Thank you that is the first answer on this that makes."
  },
  "source": "meta"
}