{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70203,"databundleVersionId":8068726,"sourceType":"competition"}],"dockerImageVersionId":30732,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Background","metadata":{}},{"cell_type":"markdown","source":"In this notebook I'll recap the many lessons I learned while competing in the BirdCLEF 2024 competition. \n\nI started competing in live Kaggle competitions this year and my goal for the year was to finish a live competition in the top 50% of the final private leaderboard. I'm super happy to say that I achieved that goal!! I finished in the top 34% of the BirdCLEF 2024 competition, ranking 329 out 991 teams. My final Private Score was **0.61**. First place scored _0.69_ and last place scored _0.46_.","metadata":{}},{"cell_type":"markdown","source":"## Results","metadata":{}},{"cell_type":"markdown","source":"I submitted 268 submissions over the course of this competition. I started by training most of the \"top-15\" models listed in Jeremy Howard's analysis [The best vision models for fine-tuning](https://www.kaggle.com/code/jhoward/the-best-vision-models-for-fine-tuning) and then picking the top five architectures to use the rest of the way. I chose 5 architectures so that each day I could submit a model from each architecture.\n\nHere are my top 10 submissions. The bolded submission (0.61) is the one of two I chose for the final leaderboard (the other scored 0.60).\n\nI trained all 10 top models (and 260/268 models overall) using the Free-A4000 16GB GPU with Paperspace Pro, which has a 6 hour auto-shutdown time limit.","metadata":{}},{"cell_type":"markdown","source":"_Top 10 submissions:_\n\n|Arch/Model Name|`item_tfms`|`batch_tfms`|Data|Weighted Loss|Epochs|Public|Private|\n|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|\n|convnext_tiny_in22k_CC|None|`PowerToDB(n_mels=286, n_fft=2048)`|100%|N|7|0.62|0.62|\n|convnext_tiny_in22k_BG|`Roll`|`PowerToDB(n_mels=286, n_fft=2048)`|20%|N|12|0.59|0.62|\n|convnext_tiny_in22k_AF|None|`PitchShift, PowerToDB(n_mels=128), ResizeTensorMelSpec(size=450)`|3%|N|20|0.63|0.62|\n|convnext_tiny_in22k_CK|None|`PowerToDB(n_mels=128, n_fft=512), FrequencyMasking`|100%|Y|6|0.62|0.61|\n|convnext_tiny_in22k_BZ|None|`PowerToDB(n_mels=128, n_fft=512), FrequencyMasking`|100%|N|7|0.64|**0.61**|\n|convnext_tiny_in22k_BP|None|`PowerToDB(n_mels=286, n_fft=2048), FrequencyMasking`|20%|N|12|0.60|0.61|\n|convnext_tiny_in22k_BK|`Noise`|`PowerToDB(n_mels=286, n_fft=2048)`|20%|N|12|0.60|0.61|\n|convnext_tiny_in22k_BI|`PeakNorm`|`PowerToDB(n_mels=286, n_fft=2048)`|20%|N|12|0.59|0.61|\n|convnext_tiny_in22k_BE|None|`PowerToDB(n_mels=286, n_fft=2048)`|20%|N|12|0.58|0.61|\n|convnext_tiny_in22k_BD|None|`PowerToDB(n_mels=286, n_fft=2048)`|20%|N|12|0.59|0.61|","metadata":{}},{"cell_type":"markdown","source":"A few notes on the values listed in the table:\n\nIn the **Data** column, I list the percentage of the total available competition training data that I used in that training run. Kaggle provided 24459 audio `ogg` files of varying length unevenly distributed across 182 species of birds. The following table provides details on the different datasets listed in my top 10 submissions:\n\n|Data|Number of 5-second Samples|\n|:-:|:-:|\n|100%|217814|\n|20%|44650|\n|3%|6158|\n\nIt's incredible to me that a model trained for 20 epochs on just _3%_ of the full dataset was my third best model (by Private Score). Wow!!\n\nIn the `item_tfms` and `batch_tfms` columns I have listed the item and batch transforms that I used in the fastai `DataBlock`. Some of the transforms, like `Noise`, `PeakNorm`, `Roll`, `PitchShift`, and `FrequencyMasking` are the built-in transforms provided in the brilliant [fastxtend library](https://fastxtend.benjaminwarner.dev/audio.03_augment.html). The other transforms are one that I custom wrote for this competition (`PowerToDB` and `ResizeTensorMelSpec`). Two \"data processing\" transforms that I haven't listed (since they aren't augmentations) are fastxtend's [`AudioBlock`](https://fastxtend.benjaminwarner.dev/audio.02_data.html#audioblock) and a custom `LoadTensorAudio` transform I wrote to load `.pt` tensor files as the fastxtend [`TensorAudio`](https://fastxtend.benjaminwarner.dev/audio.01_core.html#tensoraudio) class.\n\nFor the \"convnext_tiny_in22k_CK\" model, I used a weighted loss using the following code, which I took from the [first place submission from last year](https://www.kaggle.com/competitions/birdclef-2023/discussion/412808#:~:text=SUPER%20IMPORTANT%3A%20Class%20sampling%20weights):\n\n```python\nts = pd.DataFrame({'paths': get_files('train_chunks_full')})\nts['label'] = ts['paths'].apply(lambda x: Path(x).parent.name)\nlabel_counts = ts.groupby('label')['paths'].count()\n\nweights = (label_counts / len(ts)) ** (-0.5)\nt_weights = torch.tensor(weights.values).float()\n```\n\nIn other words, I took 1 over the square root of each bird species class' proportion in the dataset as its weight.","metadata":{}},{"cell_type":"markdown","source":"## Lessons Learned","metadata":{}},{"cell_type":"markdown","source":"### Resiliency and Consistency","metadata":{}},{"cell_type":"markdown","source":"I consider myself an \"advanced beginner\" in machine learning so I'm facing a skill issue when it comes to these competitions. There are just certain things I do not know, and that translates to a lack of intuition on what will and won't work. Only time and experience will build that intuition. \n\nSo, what I focused on instead was to practice the skills of resiliency and consistency. I made sure to submit 5 submissions every day (consistency) which was really hard to do during a stretch of 33 days where my submissions didn't equal or improve my Public score, but I pushed through regardless (resiliency). Those two skills are transferable to my future self who will be more experienced and knowledgable than I am today. I kept reminding myself that one day I will have the intuition and skill to build better models.","metadata":{}},{"cell_type":"markdown","source":"### Using the fastxtend Library","metadata":{}},{"cell_type":"markdown","source":"This was the first time I have worked with audio data and the first time I have used the wonderful [fasxtend library](https://fastxtend.benjaminwarner.dev/). To get familiarized with the library, I wrote a notebook [BirdCLEF2024 Getting Started with fastxtend Audio](https://www.kaggle.com/code/vishalbakshi/birdclef2024-getting-started-with-fastxtend-audio) in which I walk through an audio classification training using the Environmental Sound Classification 50 dataset.\n\nSomething Jeremy teaches us in the [Practical Deep Learning for Coders course](https://course.fast.ai/) is to always check that your `DataBlock` works as expected using the `DataBlock.summary` method. I used this before every one of my 268 training runs. During one of these runs, I came across the following error when the model tried to apply `PitchShift` (or any `BatchRandTransform`) to the batch:\n\n```python\nTypeError: __call__() missing 1 required positional argument: 'split_idx'\n```\nBasically, the `split_idx` value was not getting passed to the transform's `before_call` method.\n\nI opened up [a GitHub issue](https://github.com/warner-benjamin/fastxtend/issues/23) in the library's repo and Benjamin Warner, the author and maintainer of this library, provided a fix:\n\n```python\nfrom fastxtend.transform import BatchRandTransform\n\ndef call_fix(self,\n    b:Tensor|tuple[Tensor,...], # Batch item(s)\n    split_idx:int|None = None, # Train (0) or valid (1) index\n    **kwargs\n) -> Tensor|tuple[Tensor,...]:\n    \"Call `super().__call__` if `self.do`\"\n    self.before_call(b, split_idx=split_idx)\n    return super().__call__(b, split_idx=split_idx, **kwargs) if self.do else b\n    \nBatchRandTransform.__call__ = call_fix\n```\nI ran that chunk of code before each training the rest of the way and was able to use the transforms successfully.","metadata":{}},{"cell_type":"markdown","source":"### Running Inference with XML Model Format","metadata":{}},{"cell_type":"markdown","source":"The most defining and challenging characteristic of this competition was its 2 hour inference time limit. Thankfully [this notebook](https://www.kaggle.com/code/honglihang/openvino-is-all-you-need) provided a nifty solution to this problem---exporting your model to fp16 XML format.\n\nI ran a little experiment in the notebook [BirdCLEF24 - pth/ONNX/XML Inference Speed Analysis](https://www.kaggle.com/code/vishalbakshi/birdclef24-pth-onnx-xml-inference-speed-analysis) in which I compared runtime and prediction value consistency between the original `.pth` model and the ONNX, fp32 XML and fp16 XML exports. I found that for larger models, there was considerable information loss when exporting to fp16 XML. Regardless, without this solution I would not have been able to participate in this competition and it was the biggest hurdle I had to overcome.\n\nSince inference is run on a notebook without internet connection, I had to rely on uploaded datasets with wheel files for the Python libraries I needed. Many thanks to the creators of the following datasets that I used for timm, openvino and ONNX related packages:\n\n- https://www.kaggle.com/datasets/kashiwaba/timm-0613\n- https://www.kaggle.com/datasets/nohamagdy/pakeges\n- https://www.kaggle.com/datasets/zijiangyang1116/onnxruntime\n- https://www.kaggle.com/datasets/ludovick/onnxruntime","metadata":{}},{"cell_type":"markdown","source":"### Creating Custom fastai `Transform`s","metadata":{}},{"cell_type":"markdown","source":"This was the first time I created non-trivial custom `Transform`s to use in my fastai `DataBlock`. The most exciting one, which took me through a couple of rabbit holes deep into the fastxtend codebase was `PowerToDB`. \n\nI wanted to utilize the `librosa` libarary's `power_to_db` function (which takes a Mel Spectrogram and scales the values to decibels) because it seemed to provide more signal to the model. When I trained my initial models using the fastxtend `MelSpectrogram`, I got a lower Public score than using this `PowerToDB`, so I stuck with `PowerToDB` the rest of the way:\n\n\n```python\nclass PowerToDB(Transform):\n    order = 75\n    def __init__(self, sr=32000, n_fft=512, n_mels=224): \n        self.sr = sr\n        self.n_fft=n_fft\n        self.n_mels=n_mels\n        self.mel = MelSpectrogram(sample_rate=self.sr, n_fft=self.n_fft, n_mels=self.n_mels)\n    def encodes(self, x:TensorAudio):\n        return TensorMelSpec.create(librosa.power_to_db(self.mel(x).cpu()), settings={})\n    \n    def to(self, *args, **kwargs):\n        device, dtype, non_blocking, convert_to_format = torch._C._nn._parse_to(*args, **kwargs)\n        self.mel.to(device)\n```\n\nA couple of things I just had to hack together and hope they worked (which they did!) was passing the empty dictionary to the `settings` parameter of `TensorMelSpec.create` and the `to` method which transferred the `MelSpectogram` to the correct device.\n\nI also wrote the following custom transforms:\n\n`LoadTensorAudio` which takes a filename to a `.pt` file (and a sample rate) and loads it using `torch.load` and then passes it to `TensorAudio` to create that object:\n\n```python\nclass LoadTensorAudio(Transform):\n    def __init__(self, sr=32000):\n        self.sr = sr\n    def encodes(self, x:Path):\n        x = torch.load(x)\n        return TensorAudio(x, sr=self.sr)\n```\n\n`ResizeTensorMelSpec` which slices a `TensorMelSpec` object so that I can control the number of elements (for example, to set them to `224` when using the `swinv2` architecture):\n\n```python\nclass ResizeTensorMelSpec(Transform):\n    order = 76\n    def __init__(self, size=224):\n        self.size = size\n    def encodes(self, x:TensorMelSpec):\n        return x[...,:self.size]\n```\n\n`AudioNormalizeWrapper` which takes as input a `TensorMelSpec` object, and passes its standard deviation and mean to the fastxtend built-in `AudioNormalize` method:\n\n```python\nclass AudioNormalizeWrapper(Transform):\n    order = 75\n    def encodes(self, x:TensorMelSpec):\n        return AudioNormalize(mean=x.mean(), std=x.std())(x)\n```","metadata":{}},{"cell_type":"markdown","source":"### Sampling Data","metadata":{}},{"cell_type":"markdown","source":"In the fastai course, Jeremy encourages us to iterate quickly. This was the first competition I've participated in where the full training data was too large to use if I wanted to iterate quickly. For example, when using the full dataset, one epoch took about 45 minutes to run. \n\nI created a helper function to sample different subsets of the full dataset. I chose to sample a maximum of ~35 five-second chunks for each class which resulted in 6158 samples across 182 classes. This was about 3% of the full dataset and took about 1 minute per epoch to train.\n\nI later reused this function to sample 20% and 40% of the full dataset and use those subsets to train dozens of models:\n\n```python\ndef sample_chunks(group, sample_size):\n        \"\"\"\n        Function to sample `sample_size` chunks from each group, or include all chunks if the group has fewer than `sample_size` chunks.\n        \"\"\"\n        # If the group has fewer than `sample_size` chunks, return the entire group\n        if len(group) < sample_size:\n            return group\n\n        # Otherwise, sample `sample_size` chunks randomly without replacement\n        return group.sample(n=sample_size, replace=False)\n\ndef generate_sampled_df(sample_size, df):\n    # Convert the DataFrame columns to NumPy arrays\n    primary_labels = df['primary_label'].values\n    filenames = df['filename'].values\n    int_chunks = df['int_chunks'].values.astype(int)\n\n    # flattened array of filenames repeated for each chunk\n    # this is much faster in NumPy than in pandas\n    repeated_filenames = np.repeat(filenames, int_chunks)\n\n    # flattened array of primary_labels repeated for each chunk\n    repeated_primary_labels = np.repeat(primary_labels, int_chunks)\n    \n    # Create a new DataFrame from the flattened arrays\n    expanded_df = pd.DataFrame({\n        'primary_label': repeated_primary_labels,\n        'filename': repeated_filenames\n    })\n\n    # Group the DataFrame by 'primary_label'\n    grouped = expanded_df.groupby('primary_label')\n\n    # Apply the `sample_chunks` function to each group and concatenate the results\n    sampled_df = pd.concat([sample_chunks(group, sample_size=sample_size) for _, group in grouped])\n\n    # Reset the index of the sampled DataFrame\n    sampled_df.reset_index(drop=True, inplace=True)\n    \n    # calculate count of each filename\n    sampled_df = sampled_df.groupby(['primary_label', 'filename'])['filename'].count().reset_index(level=0).rename(columns={\"filename\": \"int_chunks\"}).reset_index()\n    \n    return sampled_df\n```\n\nI frequently utilized Claude's help in writing these functions.","metadata":{}},{"cell_type":"markdown","source":"### Disk Space and Runtime Trade-Off for `ogg` and `pt` Files","metadata":{}},{"cell_type":"markdown","source":"The last challenge/solution that I'll discuss in this recap is the trade-off between the size of file on disk and the amount of time it takes to generate the file. I chose to train my models on 5-second audio segments because the inference was to be run on 5-second audio segments of the test set.\n\nInitially, I chunked the `ogg` files (after converting them to `TensorAudio` objects) into 5-second-long tensors and saved them as `.pt` files. This took up a considerable amount of disk space. Each 5-second chunk took up 641 kB, which made it impossible to store the full dataset on disk as it would take up about 641000 bytes x 217814 5-second chunks / 1e9 bytes/GB = 140 GB of disk space (plus the 23 GB of original `ogg` audio files) and I only had 150 GB of storage max on Paperspace. However, it took only about 4 ms to save each `.pt` file which was great during inference since it only took somewhere around 5 minutes to chunk up and store the ~1100 or so test files that were each 4 minutes long.\n\nIn order to train on 5-second chunks of the full dataset, I decided to save them as `ogg` files using the `soundfile` library. Each 5-second `ogg` file took up only 38 kB of disk space, about 17 times smaller than a `.pt` file. This was great for training as I could more than easily store the full ~8GB of 217814 5-second `ogg` files using Paperspace. The trade-off here was that it took abnout 8 times as long (32 ms) to save each chunk! This was a disaster for inference and it caused two of my submissions to time out. \n\nMy solution was to utilize `fastcore`'s `parallel` function to process the chunking and saving of data with multiple CPU cores. This allowed me to use `ogg` chunks during training and inference without notebook timeout.","metadata":{}},{"cell_type":"markdown","source":"## Final Thoughts","metadata":{}},{"cell_type":"markdown","source":"On a personal note, I love birds. My African Gray passed away a few years ago and not a day goes by that I don't miss his songs and calls. Listening to the different bird sounds in this competition's dataset over the past couple of months was soothing and healing.\n\nThis competition was hard enough that I had to learn something new nearly every day, but doable enough that I could find solutions and overcome obstacles along the way. I am really proud of myself for persevering through the tough stretches when my submissions were not improving my score. I am also proud of myself for achieving my goal of finishing in the top 50% in 2024. I look forward to competing in more competitions this year.\n\nI hope you enjoyed this recap! Please upvote this notebook if you did. \n\nYou can find me on Twitter [@vishal_learner](https://twitter.com/vishal_learner).","metadata":{}}]}