{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30684,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I am sharing some of my insights while exploring the competition's\ndata.\n\nPlease leave a comment about additional things you\nwant me to check/explore further.\n\nEnjoy!\n\n\n\n![image](https://www.skyisland.in/uploads/1/2/1/7/121732579/image-19_orig.jpg)","metadata":{}},{"cell_type":"markdown","source":"# Some Spectrograms","metadata":{}},{"cell_type":"markdown","source":"We can start by reading some of the `.ogg` files and creating different spectrograms. For that, we can use the `librosa` library. \n\nhttps://librosa.org/doc/latest/index.html","metadata":{}},{"cell_type":"markdown","source":"## STFT","metadata":{}},{"cell_type":"markdown","source":"Let's start with a short-time Fourier transform plot (STFT)","metadata":{}},{"cell_type":"code","source":"import librosa\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\ndef stft_plot(path):\n    y, sr = librosa.load(path)\n    target = Path(path).parent.stem\n    name = Path(path).stem\n\n    D = librosa.stft(y)\n\n    magnitude = np.abs(D)\n    phase = np.angle(D)\n\n    plt.figure(figsize=(10, 4))\n    librosa.display.specshow(librosa.amplitude_to_db(magnitude, ref=np.max), y_axis='log', x_axis='time')\n    plt.colorbar(format='%+2.0f dB')\n    plt.title(f'STFT Magnitude Spectrogram for {target}/{name}')\n    plt.show()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-06-04T19:39:25.722430Z","iopub.execute_input":"2024-06-04T19:39:25.723296Z","iopub.status.idle":"2024-06-04T19:39:25.765535Z","shell.execute_reply.started":"2024-06-04T19:39:25.723259Z","shell.execute_reply":"2024-06-04T19:39:25.764699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = \"../input/birdclef-2024/train_audio/asbfly/XC134896.ogg\"\nstft_plot(path)","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:39:25.767372Z","iopub.execute_input":"2024-06-04T19:39:25.767702Z","iopub.status.idle":"2024-06-04T19:39:39.700548Z","shell.execute_reply.started":"2024-06-04T19:39:25.767667Z","shell.execute_reply":"2024-06-04T19:39:39.699342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can loop over multiple files: ","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\ntarget = \"zitcis1\"\nbase_path = \"../input/birdclef-2024/train_audio/{target}\"\nto_yield = 5\npaths = Path(base_path.format(target=target)).glob(\"*.ogg\")\nfor _ in range(to_yield):\n    path = next(paths)\n    stft_plot(path)","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:39:39.701855Z","iopub.execute_input":"2024-06-04T19:39:39.702307Z","iopub.status.idle":"2024-06-04T19:39:49.753842Z","shell.execute_reply.started":"2024-06-04T19:39:39.702278Z","shell.execute_reply":"2024-06-04T19:39:49.752792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Other transformations?","metadata":{}},{"cell_type":"markdown","source":"Let's see if we can make other plots based on signal transformations:\n\n* MFCCS: short for Mel-Frequency Cepstral Coefficients\n* CQT: short for constant-q transform\n* Wavelet transform: this is based on wavelets","metadata":{}},{"cell_type":"markdown","source":"## MFCC","metadata":{}},{"cell_type":"code","source":"import librosa\nimport librosa.display\nimport matplotlib.pyplot as plt\n\n\ndef mfcc_plot(path):\n    target = Path(path).parent.stem\n    name = Path(path).stem\n    y, sr = librosa.load(path)\n    mfccs = librosa.feature.mfcc(y=y, sr=sr)\n    librosa.display.specshow(mfccs, sr=sr, x_axis='time')\n    plt.colorbar()\n    plt.title(f'MFCC for {target}/{name}')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:39:49.756314Z","iopub.execute_input":"2024-06-04T19:39:49.757116Z","iopub.status.idle":"2024-06-04T19:39:49.764234Z","shell.execute_reply.started":"2024-06-04T19:39:49.757080Z","shell.execute_reply":"2024-06-04T19:39:49.763228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\ntarget = \"zitcis1\"\nbase_path = \"../input/birdclef-2024/train_audio/{target}\"\nto_yield = 5\npaths = Path(base_path.format(target=target)).glob(\"*.ogg\")\nfor _ in range(to_yield):\n    path = next(paths)\n    mfcc_plot(path)","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:39:49.765702Z","iopub.execute_input":"2024-06-04T19:39:49.766487Z","iopub.status.idle":"2024-06-04T19:39:53.453842Z","shell.execute_reply.started":"2024-06-04T19:39:49.766425Z","shell.execute_reply":"2024-06-04T19:39:53.452727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## CQT","metadata":{}},{"cell_type":"markdown","source":"Let's finish this part with the CQT.","metadata":{}},{"cell_type":"code","source":"def cqt_plot(path):\n    target = Path(path).parent.stem\n    name = Path(path).stem\n    y, sr = librosa.load(path)\n    cqt = librosa.cqt(y, sr=sr)\n    librosa.display.specshow(librosa.amplitude_to_db(np.abs(cqt), ref=np.max), sr=sr, x_axis='time', y_axis='cqt_note')\n    plt.colorbar(format='%+2.0f dB')\n    plt.title(f'CQT for {target}/{name}')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:39:53.455273Z","iopub.execute_input":"2024-06-04T19:39:53.455694Z","iopub.status.idle":"2024-06-04T19:39:53.462550Z","shell.execute_reply.started":"2024-06-04T19:39:53.455639Z","shell.execute_reply":"2024-06-04T19:39:53.461300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target = \"zitcis1\"\nbase_path = \"../input/birdclef-2024/train_audio/{target}\"\nto_yield = 5\npaths = Path(base_path.format(target=target)).glob(\"*.ogg\")\nfor _ in range(to_yield):\n    path = next(paths)\n    cqt_plot(path)","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:39:53.464195Z","iopub.execute_input":"2024-06-04T19:39:53.464531Z","iopub.status.idle":"2024-06-04T19:39:59.664346Z","shell.execute_reply.started":"2024-06-04T19:39:53.464503Z","shell.execute_reply":"2024-06-04T19:39:59.663171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore the metadata","metadata":{}},{"cell_type":"markdown","source":"Let's explore some of the metadata from the `train_metadata.csv` file.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ntrain_path = \"../input/birdclef-2024/train_metadata.csv\"\ndf = pd.read_csv(train_path)","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:39:59.665916Z","iopub.execute_input":"2024-06-04T19:39:59.666346Z","iopub.status.idle":"2024-06-04T19:40:01.006407Z","shell.execute_reply.started":"2024-06-04T19:39:59.666309Z","shell.execute_reply":"2024-06-04T19:40:01.005367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.sample(5)","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:40:01.008769Z","iopub.execute_input":"2024-06-04T19:40:01.009105Z","iopub.status.idle":"2024-06-04T19:40:01.037038Z","shell.execute_reply.started":"2024-06-04T19:40:01.009078Z","shell.execute_reply":"2024-06-04T19:40:01.036021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[\"primary_label\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:40:01.038264Z","iopub.execute_input":"2024-06-04T19:40:01.038584Z","iopub.status.idle":"2024-06-04T19:40:01.053802Z","shell.execute_reply.started":"2024-06-04T19:40:01.038556Z","shell.execute_reply":"2024-06-04T19:40:01.052713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems there are **182** primary labels (i.e. 182 species to \npredict).\nSome of the species have multiple audio files (500), while\nothers have only a few (5).\n\nThis distribution looks a bit similar to the one from last year's competition (and there is most likely a [bug](https://www.kaggle.com/competitions/birdclef-2024/discussion/490990#2771924) in some targets)\n\nHere is a barplot of the most popular species (in\nterms of audio files): ","metadata":{}},{"cell_type":"code","source":"s = df[\"primary_label\"].value_counts()\ns[s > 250].plot(kind=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2024-06-04T19:40:01.055231Z","iopub.execute_input":"2024-06-04T19:40:01.055671Z","iopub.status.idle":"2024-06-04T19:40:01.530351Z","shell.execute_reply.started":"2024-06-04T19:40:01.055615Z","shell.execute_reply":"2024-06-04T19:40:01.529260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Metric","metadata":{}},{"cell_type":"markdown","source":"It is a custom [ROC-AUC](https://en.wikipedia.org/wiki/Receiver_operating_characteristic) with macro average\nover classes that have true positive rate.\n\nAn implementation can be found here: https://www.kaggle.com/code/metric/birdclef-roc-auc\n\nLet's try to understand how this metric works given some examples.","metadata":{}},{"cell_type":"markdown","source":"The ROC-AUC metric is computed by first computing two sub-metrics:\n    \n- TPR: true positive rate, also known as **recall** (R). This is the ratio of correctly predicted positives over all positives, i.e. $TPR:=\\frac{TP}{TP+FN}$\n- FPR: false positive rate. This the ratio of incorrectly predicted positives (i.e. there should have been predcited negatives) over all negatives, i.e. $FPR:=\\frac{FP}{TN+FP}$\n\nOnce we have these two rates, we can vary the classification threshold (i.e. how \nto decide if the probabilty belongs to the positive or negative class) and thus get a curve:\n\n![roc](https://upload.wikimedia.org/wikipedia/commons/1/13/Roc_curve.svg)\n\n","metadata":{}},{"cell_type":"markdown","source":"Now, for this competition, if the TP for one class is 0 (i.e. all the labels are 0 for this class), it isn't used to compute the macro ROC-AUC score: there is a first filtering step and then the [sklearn.metrics.roc_auc_score](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html) is used.","metadata":{}},{"cell_type":"markdown","source":"Some additional useful terminology:\n* Predicted positives: $PP := TP+FP = FN + FP = FN + TN$\n* Predicted negatives: $PN := TN+FN $\n* Positives: $P := TP + FN$\n* Negatives: $N := TN + FP$","metadata":{}},{"cell_type":"markdown","source":"# Past solutions","metadata":{}},{"cell_type":"markdown","source":"Let's explore last year's top solutions, especially the first place: https://www.kaggle.com/competitions/birdclef-2023/discussion/412808","metadata":{}},{"cell_type":"markdown","source":"We can try a [SED](https://paperswithcode.com/task/sound-event-detection) (sound event detection) model with the following backbone to get started: \n[eca_nfnet](https://huggingface.co/timm/eca_nfnet_l0).","metadata":{}},{"cell_type":"markdown","source":"## CV splitting","metadata":{}},{"cell_type":"markdown","source":"As done in previous years, we can try a **stratified CV on 5 Folds**. \nThe stratification is done over the distribution of the targets.\nWe can also try to group using the author of each sound clip.","metadata":{}},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"markdown","source":"Notice that for this competition, **GPU compute is disabled** and you\nwill have to think about adapting your code to CPU usage (you have **120 minutes** to score the test files).\n\nThis inference notebook from last year's competition gives you some tips on how to do it on time: https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook","metadata":{}},{"cell_type":"markdown","source":"# Keywords and useful links","metadata":{}},{"cell_type":"markdown","source":"As usual, it is always useful to have some keywords/concepts and links\nto dig deeper and understand the subject.","metadata":{}},{"cell_type":"markdown","source":"- The thing we are exploring: Indian bird species.\n- Website: https://stateofindiasbirds.in/\n- Another website: https://www.skyisland.in/\n- Location: western ghats https://en.wikipedia.org/wiki/Western_Ghats\n- PAM: passive acoustic monitoring.\n- Past top solutions of the 2023 competition: https://www.kaggle.com/competitions/birdclef-2024/discussion/490895#2771413\n- API to interact with the website https://xeno-canto.org/: https://github.com/ntivirikin/xeno-canto-py\n- This website can be useful to explore each species: https://ebird.org/species\n- First place of 2023 Github repo: https://github.com/VSydorskyy/BirdCLEF_2023_1st_place\n- Feed a 1D tensor into a model that expects 3D tensor (CV model for example): https://timm.fast.ai/models#Case-1:-When-the-number-of-input-channels-is-1\n- Inference notebook that can be used as an inspiration: https://www.kaggle.com/code/vladimirsydor/bird-clef-2023-inference-v1/notebook\n- Additional data here: https://www.kaggle.com/competitions/birdclef-2024/discussion/491687\n- A very tourough audio tutorial from pytorch audio: https://pytorch.org/audio/main/tutorials/audio_feature_extractions_tutorial.html#sphx-glr-tutorials-audio-feature-extractions-tutorial-py\n- Google bird vocalization classifier: https://www.kaggle.com/models/google/bird-vocalization-classifier","metadata":{"execution":{"iopub.status.busy":"2024-04-16T12:20:53.864062Z","iopub.execute_input":"2024-04-16T12:20:53.864475Z","iopub.status.idle":"2024-04-16T12:20:53.906859Z","shell.execute_reply.started":"2024-04-16T12:20:53.864442Z","shell.execute_reply":"2024-04-16T12:20:53.905292Z"}}}]}