{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":19596,"databundleVersionId":1292430,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# Load libraries\nimport os\nimport warnings\nimport librosa\nimport numpy as np\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import LabelEncoder\nfrom tqdm import tqdm\n\n# Suppress all warnings\nwarnings.filterwarnings(\"ignore\")\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-04T06:06:25.072455Z","iopub.execute_input":"2024-12-04T06:06:25.072961Z","iopub.status.idle":"2024-12-04T06:06:25.080347Z","shell.execute_reply.started":"2024-12-04T06:06:25.072918Z","shell.execute_reply":"2024-12-04T06:06:25.078681Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Processing\n\nMost of the data has not undergone extensive processing, including proper handling of metadata. This may result in warnings, such as \"Xing stream size off by more than 1%,\" leading to increased fuzziness during operations like seeking. For improved reliability, the more processed MP3 files will serve as the base data, which will then be converted to WAV format for comprehensive model development. WAV files are ideal for this purpose as they store uncompressed, lossless audio data.","metadata":{}},{"cell_type":"code","source":"class BirdDataset(Dataset):\n    def __init__(self, file_paths, labels, sr=32000, duration=5, n_mels=64, max_frames=313):\n        self.file_paths = file_paths\n        self.labels = labels\n        self.sr = sr\n        self.duration = duration\n        self.n_mels = n_mels\n        self.max_frames = max_frames\n\n    def __len__(self):\n        return len(self.file_paths)\n\n    def __getitem__(self, idx):\n        file_path = self.file_paths[idx]\n        label = self.labels[idx]\n\n        # Load audio and compute mel spectrogram\n        waveform, _ = librosa.load(file_path, sr=self.sr, duration=self.duration, mono=True)\n        mel_spec = librosa.feature.melspectrogram(y=waveform, sr=self.sr, n_mels=self.n_mels)\n        mel_spec = librosa.power_to_db(mel_spec, ref=np.max)\n\n        # Normalize and pad/truncate spectrogram to fixed size\n        mel_spec = (mel_spec - mel_spec.mean()) / mel_spec.std()\n        if mel_spec.shape[1] < self.max_frames:\n            pad_width = self.max_frames - mel_spec.shape[1]\n            mel_spec = np.pad(mel_spec, ((0, 0), (0, pad_width)), mode=\"constant\")\n        else:\n            mel_spec = mel_spec[:, :self.max_frames]\n\n        mel_spec = torch.tensor(mel_spec, dtype=torch.float32).unsqueeze(0)  # Add channel dimension\n        return mel_spec, label\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T06:06:25.082488Z","iopub.execute_input":"2024-12-04T06:06:25.082883Z","iopub.status.idle":"2024-12-04T06:06:25.103734Z","shell.execute_reply.started":"2024-12-04T06:06:25.082838Z","shell.execute_reply":"2024-12-04T06:06:25.102177Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## CNN Model Architecture: **BirdClassifier**\n\nThe `BirdClassifier` is a convolutional neural network (CNN) designed for audio classification, specifically to classify bird species based on their sound spectrograms. Here's a breakdown of its architecture:\n\n### **Initialization (`__init__`)**\n1. **Input Parameters:**\n   - `num_classes`: Number of bird species to classify.\n   - `n_mels`: Number of Mel-frequency bins in the spectrogram (default is 64).\n   - `max_frames`: Maximum time frames in the spectrogram (default is 313).\n\n2. **Convolutional Blocks:**\n   - **`conv_block1`**:\n     - First convolution layer with 32 filters, kernel size 3x3, stride 1, and padding 1.\n     - Activation: ReLU.\n     - Max pooling with a 2x2 kernel to reduce spatial dimensions by half.\n   - **`conv_block2`**:\n     - Second convolution layer with 64 filters, same kernel and pooling parameters.\n   - **`conv_block3`**:\n     - Third convolution layer with 128 filters, same kernel and pooling parameters.\n   - Each block progressively extracts more complex features from the input data while reducing its spatial dimensions.\n\n3. **Fully Connected Layers:**\n   - **`fc1`**:\n     - A fully connected layer with 256 hidden units that connects the flattened output of the last convolutional block to a dense representation.\n   - **`fc2`**:\n     - Final fully connected layer that maps the dense representation to the number of classes (`num_classes`).\n\n4. **Output Size Calculation:**\n   - After three convolutional blocks, the spatial dimensions of the spectrogram are reduced by a factor of \\(2^3 = 8\\) (due to max pooling).\n   - `conv_output_size = (n_mels // 8) * (max_frames // 8)`.\n   - The feature map is then flattened and passed to the fully connected layers.\n\n---\n\n### **Forward Method**\n1. **Input:**\n   - A spectrogram-like tensor `x` with shape `(batch_size, 1, n_mels, max_frames)`.\n\n2. **Process Flow:**\n   - Pass `x` through `conv_block1`, `conv_block2`, and `conv_block3` sequentially.\n   - Flatten the feature map into a 1D vector for each batch using `x.flatten(1)`.\n   - Pass the flattened vector through the fully connected layers (`fc1` and `fc2`).\n\n3. **Output:**\n   - Returns logits (unactivated scores) for each class, which can be passed through a softmax function to produce class probabilities.\n\n---\n\n### **Summary of Architecture**\n- **Input**: Spectrogram (1 channel, `n_mels` x `max_frames`).\n- **Feature Extraction**: Three convolutional blocks with increasing filters (32, 64, 128) and ReLU activation.\n- **Dimensionality Reduction**: Max pooling after each convolutional block.\n- **Classification**: Two fully connected layers to map extracted features to class logits.\n- **Output**: Predictions for `num_classes`.\n","metadata":{}},{"cell_type":"code","source":"class BirdClassifier(nn.Module):\n    def __init__(self, num_classes, n_mels=64, max_frames=313):\n        super(BirdClassifier, self).__init__()\n        self.conv_block1 = nn.Sequential(\n            nn.Conv2d(1, 32, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(kernel_size=2)\n        )\n        self.conv_block2 = nn.Sequential(\n            nn.Conv2d(32, 64, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(kernel_size=2)\n        )\n        self.conv_block3 = nn.Sequential(\n            nn.Conv2d(64, 128, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(kernel_size=2)\n        )\n        conv_output_size = (n_mels // 8) * (max_frames // 8)\n        self.fc1 = nn.Linear(128 * conv_output_size, 256)\n        self.fc2 = nn.Linear(256, num_classes)\n\n    def forward(self, x):\n        x = self.conv_block1(x)\n        x = self.conv_block2(x)\n        x = self.conv_block3(x)\n        x = x.flatten(1)\n        x = self.fc1(x)\n        x = self.fc2(x)\n        return x\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T06:06:25.105115Z","iopub.execute_input":"2024-12-04T06:06:25.105507Z","iopub.status.idle":"2024-12-04T06:06:25.122260Z","shell.execute_reply.started":"2024-12-04T06:06:25.105470Z","shell.execute_reply":"2024-12-04T06:06:25.120896Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Functions Summary\n\n### **`train_model` Function**\n- **Purpose**: Trains the model on the given dataset.\n- **Steps**:\n  1. Sets the model to training mode.\n  2. Iterates through training batches from the DataLoader.\n  3. Moves inputs and labels to the specified device (CPU/GPU).\n  4. Computes predictions, calculates the loss, and backpropagates the gradients.\n  5. Updates model parameters using the optimizer.\n  6. Accumulates the loss for tracking.\n- **Output**: Returns the average loss across all training batches.\n\n### **`evaluate_model` Function**\n- **Purpose**: Evaluates the model's performance on the validation/test dataset.\n- **Steps**:\n  1. Sets the model to evaluation mode.\n  2. Disables gradient computation for efficiency.\n  3. Iterates through validation batches.\n  4. Moves inputs and labels to the specified device.\n  5. Computes predictions and compares them with true labels.\n  6. Tracks the total and correctly predicted samples.\n- **Output**: Returns the accuracy as the ratio of correct predictions to total samples.\n","metadata":{}},{"cell_type":"code","source":"def train_model(model, train_loader, criterion, optimizer, device):\n    model.train()\n    running_loss = 0.0\n    for inputs, labels in tqdm(train_loader, desc=\"Training Batches\", leave=False):\n        inputs, labels = inputs.to(device), labels.to(device)\n        optimizer.zero_grad()\n        outputs = model(inputs)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n        running_loss += loss.item()\n    return running_loss / len(train_loader)\n\n\ndef evaluate_model(model, val_loader, device):\n    model.eval()\n    correct = 0\n    total = 0\n    with torch.no_grad():\n        for inputs, labels in tqdm(val_loader, desc=\"Evaluating Batches\", leave=False):\n            inputs, labels = inputs.to(device), labels.to(device)\n            outputs = model(inputs)\n            _, predicted = torch.max(outputs, 1)\n            total += labels.size(0)\n            correct += (predicted == labels).sum().item()\n    return correct / total\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T06:06:25.124065Z","iopub.execute_input":"2024-12-04T06:06:25.124612Z","iopub.status.idle":"2024-12-04T06:06:25.146037Z","shell.execute_reply.started":"2024-12-04T06:06:25.124558Z","shell.execute_reply":"2024-12-04T06:06:25.144801Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Filter Function for Valid Files\n\nDue to limited preprocessing of the training data, the model may encounter warnings with unsuitable files. This function acts as a preliminary filter to block invalid data. However, it is a basic implementation and not suitable for a fully developed model.","metadata":{}},{"cell_type":"code","source":"def filter_files(file_paths, labels, sr=32000, duration=5):\n    valid_file_paths = []\n    valid_labels = []\n    skipped_files = []\n\n    print(\"Validating audio files...\")\n    for i, file_path in enumerate(tqdm(file_paths, desc=\"Filtering Invalid Files\")):\n        try:\n            librosa.load(file_path, sr=sr, duration=duration, mono=True)\n            valid_file_paths.append(file_path)\n            valid_labels.append(labels[i])\n        except Exception as e:\n            skipped_files.append(file_path)\n    \n    if skipped_files:\n        print(f\"Skipped {len(skipped_files)} files due to errors.\")\n        with open(\"skipped_files.log\", \"w\") as log_file:\n            for file in skipped_files:\n                log_file.write(f\"{file}\\n\")\n    \n    return valid_file_paths, valid_labels\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T06:06:25.148405Z","iopub.execute_input":"2024-12-04T06:06:25.148793Z","iopub.status.idle":"2024-12-04T06:06:25.169195Z","shell.execute_reply.started":"2024-12-04T06:06:25.148753Z","shell.execute_reply":"2024-12-04T06:06:25.167813Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Main Function Summary\n\n1. **File Paths and Labels:**\n   - Loads audio files from the training directory.\n   - Extracts file paths and corresponding species labels for each `.mp3` file.\n\n2. **Label Encoding:**\n   - Encodes species labels into numerical format using a `LabelEncoder`.\n\n3. **File Filtering:**\n   - Filters out invalid audio files using the `filter_files` function.\n\n4. **Data Splitting:**\n   - Splits the valid data into training and testing sets with an 80/20 split.\n\n5. **Dataset and DataLoader:**\n   - Creates `BirdDataset` objects for training and testing data.\n   - Initializes DataLoaders for batch processing during training and testing.\n\n6. **Model Setup:**\n   - Initializes the `BirdClassifier` model with the number of species classes.\n   - Moves the model to the appropriate device (GPU if available, otherwise CPU).\n   - Defines the loss function (`CrossEntropyLoss`) and optimizer (`Adam`).\n\n7. **Training and Evaluation:**\n   - Trains the model for a specified number of epochs (1 in this case).\n   - Prints training loss and validation accuracy for each epoch.\n","metadata":{}},{"cell_type":"code","source":"def main():\n    # File paths and labels\n    train_audio_dir = \"/kaggle/input/birdsong-recognition/train_audio\"\n    file_paths = []\n    labels = []\n\n    for species in tqdm(os.listdir(train_audio_dir), desc=\"Loading Audio Files\"):\n        species_path = os.path.join(train_audio_dir, species)\n        if os.path.isdir(species_path):\n            for fname in os.listdir(species_path):\n                if fname.endswith(\".mp3\"):\n                    file_paths.append(os.path.join(species_path, fname))\n                    labels.append(species)\n\n    # Encode labels\n    label_encoder = LabelEncoder()\n    labels = label_encoder.fit_transform(labels)\n\n    # Filter files\n    valid_file_paths, valid_labels = filter_files(file_paths, labels)\n\n    # Split data into train and test sets\n    train_paths, test_paths, train_labels, test_labels = train_test_split(\n        valid_file_paths, valid_labels, test_size=0.2, random_state=42\n    )\n\n    # Dataset and DataLoader\n    train_dataset = BirdDataset(train_paths, train_labels)\n    test_dataset = BirdDataset(test_paths, test_labels)\n    train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True)\n    test_loader = DataLoader(test_dataset, batch_size=32, shuffle=False)\n\n    # Model, criterion, optimizer\n    num_classes = len(label_encoder.classes_)\n    model = BirdClassifier(num_classes=num_classes)\n    device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n    model.to(device)\n    criterion = nn.CrossEntropyLoss()\n    optimizer = optim.Adam(model.parameters(), lr=1e-3)\n\n    # Train the model\n    num_epochs = 1\n    for epoch in range(num_epochs):\n        print(f\"Epoch {epoch + 1}/{num_epochs}\")\n        train_loss = train_model(model, train_loader, criterion, optimizer, device)\n        val_accuracy = evaluate_model(model, test_loader, device)\n        print(f\"Loss: {train_loss:.4f}, Accuracy: {val_accuracy:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T06:06:25.171512Z","iopub.execute_input":"2024-12-04T06:06:25.171936Z","iopub.status.idle":"2024-12-04T06:06:25.186282Z","shell.execute_reply.started":"2024-12-04T06:06:25.171886Z","shell.execute_reply":"2024-12-04T06:06:25.184832Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#if __name__ == \"__main__\":\n#    main()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T06:06:25.187821Z","iopub.execute_input":"2024-12-04T06:06:25.188192Z","iopub.status.idle":"2024-12-04T06:06:25.208587Z","shell.execute_reply.started":"2024-12-04T06:06:25.188156Z","shell.execute_reply":"2024-12-04T06:06:25.207493Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusion\n\nFor the purpose of detecting bird species through vocalization, there are multiple machine learning approaches to explore, such as Convolutional Neural Networks (CNNs), Artificial Neural Networks (ANNs), Deep Neural Networks (DNNs), and others. However, the primary limitation in the current project is the lack of sufficient training data to develop a highly accurate custom model from scratch.\n\nGiven the constraints, leveraging pre-trained models is a more practical and effective solution. For instance, the **PANN (Pretrained Audio Neural Network)**, available at [PANNs GitHub](https://github.com/qiuqiangkong/audioset_tagging_cnn/), is an excellent choice. This model has been trained on over 2 million audio files, providing robust feature extraction capabilities across diverse audio data. It is supported by a Cornell research paper ([PANNs: Large-Scale Pretrained Audio Neural Networks for\nAudio Pattern Recognition](https://arxiv.org/abs/1912.10211)) that highlights its utility and performance in audio classification tasks, despite being developed a few years ago.\n\n#### Key Recommendations:\n1. **Adopt a Pre-Trained Model**: Use PANN or a similar pre-trained audio model to leverage extensive training on diverse datasets, ensuring higher accuracy without requiring significant computational resources or large-scale data.\n\n2. **Fine-Tune the Pre-Trained Model**: Adjust the pre-trained model to suit the specific requirements of bird species classification by fine-tuning it with the available dataset.\n\n3. **Data Augmentation**: Increase the diversity and robustness of your dataset by applying techniques like time-stretching, pitch-shifting, or adding noise to create synthetic variations of the existing audio files.\n\n4. **Evaluate Transfer Learning Options**: Explore other recent pre-trained models designed for audio classification tasks and compare their performance with PANN to identify the most suitable one.\n\n5. **Future Steps**: Invest in collecting or sourcing more labeled audio data for birds to build a larger dataset, enabling the possibility of training a custom model in the future.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}