{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [BirdCLEF23](https://www.kaggle.com/competitions/birdclef-2023) Generating Uniform No-Call Chunks\n\n[Original dataset collected from 3 other comps for BirdCLEF21](https://www.kaggle.com/datasets/christofhenkel/birdclef2021-background-noise)  \n[The resulting dataset from this notebook](https://www.kaggle.com/datasets/ollypowell/birdclef-8-sec-ogg)\n\n### Motivation\n\nThe motivation for this is to have unform chunks to mix in with my existing datasets, along with a compatible csv file for easy concatination, in the same format:\n\nI am working with 8 second chunks, but have based all my datasets on the same csv format as [this one](/kaggle/input/cleaned-training-labels-21-23-for-birdclef2023) for easy concantenation.\n\nI have saved to ogg this time, because I'm sharing my data around with colab, and my own machine.  In theory .wav will load a bit faster, but the file sizes were less practical for me.\n\n### Data sources\n- The one labeled *Train soundscapes* was provided for the 2021 competition, but I've assumed has had everything except the 'no-call' parts removed.  There are more files there, so the dataset could be increased.  [link here](https://www.kaggle.com/competitions/birdclef-2021/data?select=train_soundscapes)\n- ff1010bird_nocall has come from [DCASE 2018](https://dcase.community/challenge2018/task-bird-audio-detection) competition, and originally sourced from the [freesound project](https://freesound.org/)\n- AICrowd collection has come from [birdCLEF2020](https://www.aicrowd.com/clef_tasks/22/task_dataset_files?challenge_id=211)  (requires a login)\n\n### Output\n- A new dataset, with files in a single folder, labelled original_folder_nm_xxx.yyy, where xxx is a unique integer, and yyy is either .wav or .ogg, depending on the choice in config. \n- A csv file with filename, file path in the kaggle system, primary class 'no-call', secondary class empty\n\n### Usage\nYou could either modify the csv format to match your own plans, and add it to an existing dataset with a no-call class label.  Or just use the chunks as background noise for augmentation.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-20T05:01:26.89796Z","iopub.execute_input":"2023-03-20T05:01:26.89912Z","iopub.status.idle":"2023-03-20T05:01:26.908618Z","shell.execute_reply.started":"2023-03-20T05:01:26.899074Z","shell.execute_reply":"2023-03-20T05:01:26.906979Z"}}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport soundfile as sf\nfrom joblib import Parallel, delayed\nfrom pathlib import Path\nimport random \nimport librosa\nimport multiprocessing as mp\nfrom IPython.display import Audio\nimport shutil ","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:33:39.357415Z","iopub.execute_input":"2023-05-10T00:33:39.357901Z","iopub.status.idle":"2023-05-10T00:33:39.546967Z","shell.execute_reply.started":"2023-05-10T00:33:39.357856Z","shell.execute_reply":"2023-05-10T00:33:39.545628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Configuration for this notebook","metadata":{}},{"cell_type":"code","source":"class Config:\n    SR = 32000 # Sampling rate of all the source files\n    CHUNK_DURATION = 5  # Clips the files to this number of seconds.\n    CLASS_NAME = 'no-call'\n    FILE_TYPE = 'wav'  # Save to .wav will potentially mean faster loading, but larger files\n    NUM_OUT = 10_000  # By comparison the mean for all birds in the three comp years is 215 samples each of varying length\n    MAX_IMBALANCE = 100 # The maximum ratio of sampling from any dataset compared to the smallest one\n    NUM_WORKERS = 4 # For parallel processing\n","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:34:28.275154Z","iopub.execute_input":"2023-05-10T00:34:28.275650Z","iopub.status.idle":"2023-05-10T00:34:28.282860Z","shell.execute_reply.started":"2023-05-10T00:34:28.275604Z","shell.execute_reply":"2023-05-10T00:34:28.281592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Setup filepaths and create output folders\n\n- Write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n- You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{}},{"cell_type":"code","source":"in_fldrs = ['/kaggle/input/birdclef2021-background-noise/aicrowd2020_noise_30sec/noise_30sec',  #225 mminutes\n            '/kaggle/input/birdclef2021-background-noise/ff1010bird_nocall/nocall',  #960 minutes\n           '/kaggle/input/birdclef2021-background-noise/train_soundscapes/nocall', ]  #30 minutes (Could add more, but would need to filter for no-call first)\n\ndata_folder = Path('/kaggle/input')  # modify to suit\nin_csv = data_folder / 'cleaned-training-labels-21-23-for-birdclef2023/train_21_22_23.csv' # for header format only\n\n#out_folder = Path('/kaggle/temp') # will be lost outside current session\nout_folder = Path('/kaggle/working/') # to save and make into a dataset\ntemp_folder = Path('/kaggle/temp') # to store intermediate files\n\nout_dataset_name =  f'birdclef-{str(Config.CHUNK_DURATION)}-sec-{Config.FILE_TYPE}' \nout_csv = out_folder / 'birdclef-nocall.csv'\nout_sf_folder = out_folder / 'train_audio'\n\nos.makedirs(out_sf_folder, exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:34:29.992640Z","iopub.execute_input":"2023-05-10T00:34:29.993075Z","iopub.status.idle":"2023-05-10T00:34:30.003183Z","shell.execute_reply.started":"2023-05-10T00:34:29.993033Z","shell.execute_reply":"2023-05-10T00:34:30.001891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"format_df = pd.read_csv(in_csv)\ndf = pd.DataFrame(columns = format_df.columns)\ndf","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:34:33.506701Z","iopub.execute_input":"2023-05-10T00:34:33.507677Z","iopub.status.idle":"2023-05-10T00:34:33.872986Z","shell.execute_reply.started":"2023-05-10T00:34:33.507633Z","shell.execute_reply":"2023-05-10T00:34:33.871876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The goal here is to have roughly the same number of minutes from each of the sound source directories.  Going for simplicity over efficiency, I'm just breaking up all the sources into uniform chunks, and saving them all temporarily. The max is actually determined by the shortest source times `Config.MAX_IMBALANCE`.  eg 30 minutes x 3 = 90 so in this case we have (30 + 90 + 90)*60/8 = 1575 8-second files, minus a bit of truncating.  For a total number less than that, set a lower value for `Config.NUM_OUT`.","metadata":{}},{"cell_type":"code","source":"def load_ogg(path):\n    y, sr = sf.read(path, always_2d=True)\n    y = np.mean(y, 1) # For any sterio (X, 2) arrays\n    if not np.isfinite(y).all():\n        y[np.isnan(y)] = np.zeros_like(y)\n        y[np.isinf(y)] = np.max(y)\n    return y, len(y)\n\n\ndef play_audio(file_path):\n    audio_abe, sr_abe = librosa.load(file_path)\n    return Audio(data=audio_abe, rate=sr_abe)\n\n\ndef file_to_chunks(item_tuple):\n    file, save_pth = item_tuple\n    parent_name = str(file.parent.name) + '_' + str(file.stem)\n    chunk_len = Config.CHUNK_DURATION * Config.SR\n    y, length = load_ogg(file)\n    if length < chunk_len + 1:\n        print(f'{file.name} is too short ({length//Config.SR} seconds)')\n    else:\n        chunks = [y[i:i+chunk_len] for i in range(0, length, chunk_len)]\n        if len(chunks[-1]) != chunk_len:\n            chunks = chunks[:-1]\n        for idx2, chunk in enumerate(chunks):\n            fn = f'{parent_name}_{str(idx2)}.{Config.FILE_TYPE}' # could add {str(save_pth).partition(\"_\")[-1]}  for more uniqueness\n            pth = save_pth / fn\n            sf.write(pth, chunk, Config.SR)\n    return","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:34:44.306031Z","iopub.execute_input":"2023-05-10T00:34:44.306474Z","iopub.status.idle":"2023-05-10T00:34:44.320847Z","shell.execute_reply.started":"2023-05-10T00:34:44.306436Z","shell.execute_reply":"2023-05-10T00:34:44.319108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_temps = {}\nfor idx, folder in enumerate(in_fldrs):\n    in_path_list = [p for p in Path(folder).rglob('*.ogg')]\n    print(f'Processing {len(in_path_list)} ogg files in folder {folder}')\n    temp_subfolder_nm = f'no_calls_{idx}'\n    temp_save = temp_folder / temp_subfolder_nm\n    os.makedirs(temp_save, exist_ok=True)\n    items = [(p, temp_save) for p in in_path_list]\n    \n    if __name__ == '__main__':\n        pool = mp.Pool(processes=4)\n        results = pool.map(file_to_chunks, items)\n        \n    new_files = [p for p in Path(temp_save).rglob(f'*.{Config.FILE_TYPE}')]\n    all_temps[temp_subfolder_nm] = new_files \n    print(f'Made {len(new_files)} {Config.CHUNK_DURATION}-second files from folder {idx+1}')","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:34:44.999061Z","iopub.execute_input":"2023-05-10T00:34:44.999794Z","iopub.status.idle":"2023-05-10T00:37:17.949579Z","shell.execute_reply.started":"2023-05-10T00:34:44.999747Z","shell.execute_reply":"2023-05-10T00:37:17.946649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir(temp_folder)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:37:17.954323Z","iopub.execute_input":"2023-05-10T00:37:17.956019Z","iopub.status.idle":"2023-05-10T00:37:17.970396Z","shell.execute_reply.started":"2023-05-10T00:37:17.955949Z","shell.execute_reply":"2023-05-10T00:37:17.969019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Randomly sample\nrandom.seed(2023)\nmin_length = min([len(val) for val in all_temps.values()])\n\nsampled_temps = []\nfor path_list in all_temps.values():\n    if len(path_list) > min_length * Config.MAX_IMBALANCE:\n        random_sample = random.sample(path_list, min_length * Config.MAX_IMBALANCE)\n    else:\n        random_sample = path_list\n    sampled_temps = sampled_temps + random_sample\n#Now the list contains similar numbers from the different source folders, to a max ratio of MAX_IMBALANCE\n\nrandom.shuffle(sampled_temps)\n\nif len(sampled_temps) > Config.NUM_OUT:\n    sampled_temps = sampled_temps[:Config.NUM_OUT]\n\nprint(f'There are {len(sampled_temps)} filepaths in the list to be moved')","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:37:17.972211Z","iopub.execute_input":"2023-05-10T00:37:17.973161Z","iopub.status.idle":"2023-05-10T00:37:18.012080Z","shell.execute_reply.started":"2023-05-10T00:37:17.973118Z","shell.execute_reply":"2023-05-10T00:37:18.010631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_rows= []\nfor path in sampled_temps:\n    fn = str(Path(path).name)\n    destination = out_sf_folder / fn\n    shutil.move(path, destination)\n    \n    csv_file_path = f'/kaggle/input/{out_dataset_name}/train_audio/{fn}'\n    row_data = {'primary_label': Config.CLASS_NAME, \n                'secondary_labels': [], \n                'type': [], \n                'filename':fn,\n                'filepath': csv_file_path }\n    new_rows.append(row_data)\n    \ndf = df.append(new_rows, ignore_index=True)\ndf.to_csv(out_csv)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:37:18.014591Z","iopub.execute_input":"2023-05-10T00:37:18.015768Z","iopub.status.idle":"2023-05-10T00:37:42.471697Z","shell.execute_reply.started":"2023-05-10T00:37:18.015716Z","shell.execute_reply":"2023-05-10T00:37:42.470289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'There are a total of {len(os.listdir(out_sf_folder))} sound files')\nprint(f'There are {df.shape[0]} rows in the new labels dataframe')","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:37:42.475628Z","iopub.execute_input":"2023-05-10T00:37:42.476123Z","iopub.status.idle":"2023-05-10T00:37:42.496440Z","shell.execute_reply.started":"2023-05-10T00:37:42.476076Z","shell.execute_reply":"2023-05-10T00:37:42.494851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below I'm just checking a random sample of the file paths written to the csv","metadata":{}},{"cell_type":"code","source":"rand_list = random.sample(range(1, 10000), 200)\nfor num in rand_list[:4]:\n    print(df.iloc[num]['filepath'])","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:37:42.499088Z","iopub.execute_input":"2023-05-10T00:37:42.500220Z","iopub.status.idle":"2023-05-10T00:37:42.643913Z","shell.execute_reply.started":"2023-05-10T00:37:42.500156Z","shell.execute_reply":"2023-05-10T00:37:42.642364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And to varify a few files have saved correctly and don't contain birds:","metadata":{}},{"cell_type":"code","source":"def path_from_csv(idx):\n    kaggle_parents = '/kaggle/input/' + out_dataset_name\n    #new_path = '../working/' + df.iloc[rand_list[idx]]['filepath'].lstrip(kaggle_parents) #I can't see the problem!\n    new_path = '../working' + df.iloc[rand_list[idx]]['filepath'].replace(kaggle_parents, '')\n    return new_path\n\npath_from_csv(5)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:26:43.075391Z","iopub.execute_input":"2023-05-10T00:26:43.076349Z","iopub.status.idle":"2023-05-10T00:26:43.090270Z","shell.execute_reply.started":"2023-05-10T00:26:43.076292Z","shell.execute_reply":"2023-05-10T00:26:43.088760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_audio(path_from_csv(0))","metadata":{"execution":{"iopub.status.busy":"2023-05-08T18:28:43.733193Z","iopub.execute_input":"2023-05-08T18:28:43.734146Z","iopub.status.idle":"2023-05-08T18:28:58.046767Z","shell.execute_reply.started":"2023-05-08T18:28:43.734095Z","shell.execute_reply":"2023-05-08T18:28:58.045029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_audio(path_from_csv(1))","metadata":{"execution":{"iopub.status.busy":"2023-05-08T18:28:58.048526Z","iopub.execute_input":"2023-05-08T18:28:58.050007Z","iopub.status.idle":"2023-05-08T18:28:58.070735Z","shell.execute_reply.started":"2023-05-08T18:28:58.049962Z","shell.execute_reply":"2023-05-08T18:28:58.069177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_audio(path_from_csv(20))","metadata":{"execution":{"iopub.status.busy":"2023-05-08T18:28:58.072894Z","iopub.execute_input":"2023-05-08T18:28:58.073312Z","iopub.status.idle":"2023-05-08T18:28:58.134554Z","shell.execute_reply.started":"2023-05-08T18:28:58.073273Z","shell.execute_reply":"2023-05-08T18:28:58.132699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_audio(path_from_csv(45))","metadata":{"execution":{"iopub.status.busy":"2023-05-08T18:28:58.136574Z","iopub.execute_input":"2023-05-08T18:28:58.137345Z","iopub.status.idle":"2023-05-08T18:28:58.160717Z","shell.execute_reply.started":"2023-05-08T18:28:58.137303Z","shell.execute_reply":"2023-05-08T18:28:58.159320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_audio(path_from_csv(70))","metadata":{"execution":{"iopub.status.busy":"2023-05-08T18:28:58.162306Z","iopub.execute_input":"2023-05-08T18:28:58.162681Z","iopub.status.idle":"2023-05-08T18:28:58.188434Z","shell.execute_reply.started":"2023-05-08T18:28:58.162619Z","shell.execute_reply":"2023-05-08T18:28:58.186710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_audio(path_from_csv(150))","metadata":{"execution":{"iopub.status.busy":"2023-05-08T18:28:58.191082Z","iopub.execute_input":"2023-05-08T18:28:58.191466Z","iopub.status.idle":"2023-05-08T18:28:58.213361Z","shell.execute_reply.started":"2023-05-08T18:28:58.191428Z","shell.execute_reply":"2023-05-08T18:28:58.212298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_audio(path_from_csv(190))","metadata":{"execution":{"iopub.status.busy":"2023-05-08T18:28:58.215196Z","iopub.execute_input":"2023-05-08T18:28:58.215959Z","iopub.status.idle":"2023-05-08T18:28:58.237352Z","shell.execute_reply.started":"2023-05-08T18:28:58.215919Z","shell.execute_reply":"2023-05-08T18:28:58.236194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!zip \"no_bird_train_audio_10000.zip\" \"../working/train_audio/\"","metadata":{"execution":{"iopub.status.busy":"2023-05-10T00:37:42.646214Z","iopub.execute_input":"2023-05-10T00:37:42.647031Z","iopub.status.idle":"2023-05-10T00:37:43.788660Z","shell.execute_reply.started":"2023-05-10T00:37:42.646986Z","shell.execute_reply":"2023-05-10T00:37:43.787145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}