{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":8900,"databundleVersionId":862232,"sourceType":"competition"},{"sourceId":1378,"sourceType":"modelInstanceVersion","modelInstanceId":1163,"modelId":162}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Freesound General-Purpose Audio Tagging Challenge","metadata":{}},{"cell_type":"markdown","source":"## Project Notebooks\n- This project notebook is the first of three. The notebook contains a project description and the data exploration / preparation steps. The second notebook is the training notebook, and the final notebook is the final project submission.\n\n- https://www.kaggle.com/code/dadamtreece/at-freesound-training (Training)\n- https://www.kaggle.com/code/dadamtreece/at-freesound-submission (Submission)\n","metadata":{}},{"cell_type":"markdown","source":"## Project Overview\n\n#### Throughout this project we will explore the Freesound dataset and create a model to perform Mean Average at 3 (MAP@3) for final predictions and scoring. Various methods of feature extraction were explored throughout the course of this project.\n\n#### This project aims to classify audio samples into one of 41 categories. This involves extracting features from raw audio data, addressing class imbalances using synthetic oversampling, and training a deep learning model to achieve high accuracy while minimizing overfitting. Our final goal is to create a robust model capable of generalizing well on unseen data, ensuring reliable classification performance.\n\n#### Some of the steps taken in this project include:\n\n- **Data Preparation:**\n\n  - Original Dataset: The dataset consists of 9473 audio samples, each labeled with one of 41 target classes. Features were\n    extracted from raw audio files using:\n  - VGGish embeddings (pre-trained audio feature extractor).\n  - MFCCs (Mel-Frequency Cepstral Coefficients).\n  - HuBERT embeddings (deep contextualized speech representations).\n\n  - Shape Normalization: All extracted features were reshaped and padded/cropped to ensure consistent input dimensions.\n\n  - Handling Class Imbalance: SMOTE (Synthetic Minority Oversampling Technique) was used to address class imbalance by generating\n  synthetic samples for underrepresented classes.\n\n\n-  **Training and Validation Split:** A 80/20 split was used for training and validation.\n\n\n-  **Model Architecture and Hyperparameter Tuning:**\n\n   - Model: A 1D Convolutional Neural Network (CNN) architecture was designed with the following components:\n    Conv1D layers with spatial dropout and L2 regularization to extract temporal features and reduce overfitting.\n    Max Pooling layer to reduce parameter spatial dimensions while retaining important features.\n    Dense layers with dropout and batch normalization for feature integration.\n\n   - Hyperparameter Tuning:\n    Used Keras Tuner (Random Search) to optimize hyperparameters, including the number of filters, kernel sizes, dropout rates, and\n   learning rate.\n\n\n-  **Cross-Validation:** 8-fold cross-validation was used to evaluate model performance across different splits of the data.\n\n   - Ensembling: Used best cross-validation model outputs to average predictions on different folds to improve generalization.\n\n\n-  **Optimization:**\n   - ReduceLROnPlateau was used to adjust the learning rate dynamically based on validation loss during cross-fold training.\n\n   - Early stopping was implemented to terminate training when validation performance stopped improving to save GPU training\n   recsources when training was not improving.","metadata":{}},{"cell_type":"markdown","source":"# Data Prep\n\n- During the data preparation process we will explore the data and create a feature extraction strategy to transform our data for training\n- Since the training data consists of audio files, there are several options for transforming the training data to include, extracting features from the data using librosa and converting those into a numpy array, extracting spectrograms, or extracting features using pre-trained audio classification models","metadata":{"execution":{"iopub.status.busy":"2024-11-23T19:09:58.059466Z","iopub.execute_input":"2024-11-23T19:09:58.060299Z","iopub.status.idle":"2024-11-23T19:09:58.085390Z","shell.execute_reply.started":"2024-11-23T19:09:58.060232Z","shell.execute_reply":"2024-11-23T19:09:58.084091Z"}}},{"cell_type":"markdown","source":"## Import Packages","metadata":{"execution":{"iopub.status.busy":"2024-11-23T19:09:58.059466Z","iopub.execute_input":"2024-11-23T19:09:58.060299Z","iopub.status.idle":"2024-11-23T19:09:58.085390Z","shell.execute_reply.started":"2024-11-23T19:09:58.060232Z","shell.execute_reply":"2024-11-23T19:09:58.084091Z"}}},{"cell_type":"code","source":"import warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\nimport os\nimport time\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport IPython\nimport IPython.display as ipd\nimport librosa\nimport librosa.display\nimport pickle\nimport joblib\nimport random\nimport cv2\nimport wave\nimport torch\nimport torchaudio\n\nfrom scipy.signal import wiener\nfrom sklearn.model_selection import train_test_split, KFold, ShuffleSplit\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.base import BaseEstimator, TransformerMixin\nfrom sklearn.preprocessing import LabelEncoder\n\nimport tensorflow as tf\nimport tensorflow_hub as hub\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.python.keras.utils.data_utils import Sequence\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import *\nfrom tensorflow.keras.regularizers import l2\nfrom tensorflow.keras.callbacks import EarlyStopping, ModelCheckpoint, ReduceLROnPlateau\nfrom tensorflow.keras.applications import MobileNetV2\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.optimizers import AdamW\nfrom tensorflow.keras.applications import Xception\nfrom sklearn.utils.class_weight import compute_class_weight\nfrom transformers import Wav2Vec2Processor, HubertModel","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Load Training Data","metadata":{"execution":{"iopub.status.busy":"2024-11-23T19:09:58.059466Z","iopub.execute_input":"2024-11-23T19:09:58.060299Z","iopub.status.idle":"2024-11-23T19:09:58.085390Z","shell.execute_reply.started":"2024-11-23T19:09:58.060232Z","shell.execute_reply":"2024-11-23T19:09:58.084091Z"}}},{"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/freesound-audio-tagging/train_post_competition.csv\", dtype=str)\nprint(train.shape)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Explore the unique classes\n- Here we explore the number of samples and the total number of unique classes","metadata":{}},{"cell_type":"code","source":"print(\"Number of training examples=\", train.shape[0], \"  Number of classes=\", len(train.label.unique()))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(train.label.unique())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Audio Samples by category (manually verified/not verified)\n- Within this dataset, there are a percentage of the samples where the labels have been manually verified, and other which have been labeled through another classification process. Here we will look at that distribution. Since the manually labeled data may be more \"trustworthy\" we will consider using only these samples for the training.","metadata":{"execution":{"iopub.status.busy":"2024-11-23T19:46:39.971982Z","iopub.execute_input":"2024-11-23T19:46:39.972451Z","iopub.status.idle":"2024-11-23T19:46:56.206304Z","shell.execute_reply.started":"2024-11-23T19:46:39.972414Z","shell.execute_reply":"2024-11-23T19:46:56.204393Z"}}},{"cell_type":"code","source":"train['manually_verified'] = train['manually_verified'].astype(int)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"category_group = train.groupby(['label', 'manually_verified']).size().unstack(fill_value=0)\n\nsorted_category_group = category_group.reindex(category_group.sum(axis=1).sort_values().index)\n\nplt.figure(figsize=(16, 10))\nplot = sorted_category_group.plot(kind='bar', stacked=True, title=\"Number of Audio Samples per Category\", figsize=(16,10))\nplot.set_xlabel(\"Category\")\nplot.set_ylabel(\"Number of Samples\")\nplt.legend(title='Manually Verified')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- As we can see, the number of samples in each class is imbalanced, as well as the number of samples which are manually verfified vs not. We will need to consider a strategy that considers the class imbalance","metadata":{}},{"cell_type":"code","source":"verification_counts = train['manually_verified'].value_counts()\nprint(verification_counts)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Minimum samples per category = ', min(train.label.value_counts()))\nprint('Maximum samples per category = ', max(train.label.value_counts()))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels = (train.label.value_counts() / len(train)).to_frame().sort_index().T\n\ntrain_labels","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Load Sample File\n- Here we will use librosa to explore information from a sample, and look different types of extracted spectrograms","metadata":{"execution":{"iopub.status.busy":"2024-11-23T19:47:44.394534Z","iopub.execute_input":"2024-11-23T19:47:44.394970Z","iopub.status.idle":"2024-11-23T19:47:51.468265Z","shell.execute_reply.started":"2024-11-23T19:47:44.394933Z","shell.execute_reply":"2024-11-23T19:47:51.466565Z"}}},{"cell_type":"code","source":"random_row = train.sample(n=1)\nrandom_filename = random_row['fname'].values[0]\nlabel = random_row['label'].values[0]\n\naudio_dir = '/kaggle/input/freesound-audio-tagging/audio_train'\naudio_path = f\"{audio_dir}/{random_filename}\"\nsample, sr = librosa.load(audio_path, sr=None)\n\nprint(f\"Playing audio: {random_filename}\")\nipd.display(ipd.Audio(sample, rate=sr))\n\nprint(f'Length: {len(sample)/sr:.2f}s')\nprint(f'Label: {label}')\nprint(f'Sample Rate: {sr}')\n\nfig, ax = plt.subplots(4, 1, figsize=(16, 10))\n\nlibrosa.display.waveshow(sample, sr=sr, ax=ax[0])\nax[0].set(title='Temporal Signal', xlabel='Time (s)', ylabel='Amplitude')\n\n# STFT - Short Term Fourier Transform: Computes the amplitude of the frequencies for different bands over time.\nstft_result = librosa.stft(sample)\nstft_db = librosa.amplitude_to_db(abs(stft_result))\nlibrosa.display.specshow(stft_db, sr=sr, x_axis='time', y_axis='log', ax=ax[1])\nax[1].set(title='STFT Spectrogram', xlabel='Time (s)', ylabel='Frequency (Hz)')\n\n# MFCC Mel Frequency Cepstral Coefficients: Describe the instantaneous spectral envelope shape of the speech signal.\nmfccs = librosa.feature.mfcc(y=sample, sr=sr, n_mfcc=13)\nlibrosa.display.specshow(mfccs, sr=sr, x_axis='time', ax=ax[2])\nax[2].set(title='MFCC', xlabel='Time (s)', ylabel='MFCC Coefficients')\n\n#Log-Mel spectrogram: Similar to STFT but represented in the Mel scale (log transformation of the frequency scale)\nmel_spec = librosa.feature.melspectrogram(y=sample, sr=sr)\nlog_mel_spec = librosa.power_to_db(mel_spec)\nlibrosa.display.specshow(log_mel_spec, sr=sr, x_axis='time', y_axis='mel', ax=ax[3])\nax[3].set(title='Log-Mel Spectrogram', xlabel='Time (s)', ylabel='Mel Frequency (Hz)')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Create a plot to explore the length of the files based on nframe \n- We can also see by exploring the number of frames in each sample, the lengths also vary significantly. We will need to consider a strategy to make sure our samples are all a similar length when extracting features for training","metadata":{"execution":{"iopub.status.busy":"2024-11-23T19:47:44.394534Z","iopub.execute_input":"2024-11-23T19:47:44.394970Z","iopub.status.idle":"2024-11-23T19:47:51.468265Z","shell.execute_reply.started":"2024-11-23T19:47:44.394933Z","shell.execute_reply":"2024-11-23T19:47:51.466565Z"}}},{"cell_type":"code","source":"train['nframes'] = train['fname'].apply(lambda f: wave.open('../input/freesound-audio-tagging/audio_train/' + f).getnframes())\n\n_, ax = plt.subplots(figsize=(16, 4))\nsns.violinplot(ax=ax, x=\"label\", y=\"nframes\", data=train)\nplt.xticks(rotation=90)\nplt.title('Distribution of audio frames, per label', fontsize=16)\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(figsize=(16,5))\ntrain.nframes.hist(bins=100)\nplt.suptitle('Frame Length Distribution in Train and Test', ha='center', fontsize='large');","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Filter the dataframe for only manually verfified samples\n- After training multiple models, it was determined that using the manually verified data only during training yields the best results. We will now create extract only those samples for training","metadata":{}},{"cell_type":"code","source":"### Filter for manually verified data only\n\ntrain = train[train['manually_verified'] == 1]\n\ntrain.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Split into Training and Validation\n- Since we will use a split on the training data for cross-fold evaluation, here we create a split for traning and valdiation, to allow for augmentation on the training but not the validation set.\n\n- Multiple split sizes were experimented with - the 80/20 spilt yielded the best results.","metadata":{}},{"cell_type":"code","source":"train_files, test_files = train_test_split(train, test_size=0.2, random_state=42)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_files.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_files.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Training Data Augmentation (As Required)\n- Best results so far have been without data augmentation so this code is here for future exploration as required. We will not use any data augmentation for our final training features.","metadata":{}},{"cell_type":"markdown","source":"import soundfile as sf\n\naudio_dir = '/kaggle/input/freesound-audio-tagging/audio_train'\n\naudio_dir = '/kaggle/input/freesound-audio-tagging/audio_train'\naugmented_audio_dir = '/kaggle/working/augmented_audio'\n\nos.makedirs(augmented_audio_dir, exist_ok=True)\n\ndef time_stretch(y, sr):\n    return librosa.effects.time_stretch(y, rate=np.random.uniform(0.8, 1.2))\n\ndef pitch_shift(y, sr):\n    return librosa.effects.pitch_shift(y, sr=sr, n_steps=np.random.randint(-2, 3))\n\ndef add_noise(y, sr):\n    return y + 0.005 * np.random.randn(len(y))\n\ndef time_shift(y, sr):\n    return np.roll(y, shift=np.random.randint(-sr // 10, sr // 10))\n\ndef volume_scale(y, sr):\n    return y * np.random.uniform(0.7, 1.3)\n\ndef augment_audio(file_path, output_dir, file_name):\n    y, sr = librosa.load(file_path, sr=16000)\n\n    augmentations = [\n        time_stretch,\n        pitch_shift,\n        add_noise,\n        time_shift,\n        volume_scale\n    ]\n\n    augmentation = np.random.choice(augmentations)\n    y_augmented = augmentation(y, sr)\n\n    if len(y_augmented) < 5 * sr:\n        y_augmented = np.pad(y_augmented, (0, 5 * sr - len(y_augmented)))\n    else:\n        y_augmented = y_augmented[:5 * sr]\n\n    if np.max(np.abs(y_augmented)) > 0:\n        y_augmented = y_augmented / np.max(np.abs(y_augmented))\n\n    augmented_file_path = os.path.join(output_dir, file_name)\n    sf.write(augmented_file_path, y_augmented, sr)\n\n    return augmented_file_path\n\naugmented_data = []\n\nfor idx, row in train_files.iterrows():\n    file_name = row['fname']\n    file_path = os.path.join(audio_dir, file_name)\n\n    augmented_data.append({'fname': file_name, 'label': row['label'], 'path': file_path})\n\n    for aug_idx in range(2):\n        augmented_file_name = f\"{os.path.splitext(file_name)[0]}_aug_{aug_idx}.wav\"\n        augmented_file_path = augment_audio(file_path, augmented_audio_dir, augmented_file_name)\n        augmented_data.append({'fname': augmented_file_name, 'label': row['label'], 'path': augmented_file_path})\n\naugmented_train_files = pd.DataFrame(augmented_data)\n\ntrain_files = augmented_train_files\n\ntrain_audio_paths = train_files['path'].tolist()","metadata":{"execution":{"iopub.status.busy":"2024-12-03T00:56:59.181958Z","iopub.execute_input":"2024-12-03T00:56:59.182394Z","iopub.status.idle":"2024-12-03T01:00:18.566404Z","shell.execute_reply.started":"2024-12-03T00:56:59.182359Z","shell.execute_reply":"2024-12-03T01:00:18.565033Z"}}},{"cell_type":"code","source":"#train_files.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Feature Extraction\n- After exploring many different feature extraction methods (spectrogram, VGGish, HuBERT) all individually, and combining them in variations, the best results were from combining VGGish extracted features with MFCC and HuBERT features.\n- VGGish is a pre-trained audio piepleine from Google\n- HuBERT is a pre-trained audio pipeline from HuggingFace\n- Use pretrained models to extract VGGish, MFCC, and HuBERT features to use for neural network\n  * Also tried with LogMel Features, but best results are without","metadata":{}},{"cell_type":"code","source":"# Define Class and functions for feature extraction\n# When extracting features, we will also trim the dead space from each audio file using librosa, \n# and further trim each sample to a 5 second duration. After exploring different sample duration lengths in training, \n# 5 seconds was the optimal choice. For samples that are less than 5 seconds in length, they will be padded with \"whitespace\"\naudio_dir = '/kaggle/input/freesound-audio-tagging/audio_train'\n\nvggish_model = hub.load('https://kaggle.com/models/google/vggish/frameworks/TensorFlow2/variations/vggish/versions/1')\nprocessor = Wav2Vec2Processor.from_pretrained(\"facebook/hubert-large-ls960-ft\")\nhubert_model = HubertModel.from_pretrained(\"facebook/hubert-large-ls960-ft\")\n\n# VGGish feature extraction\nclass VGGishFeatureExtractor(BaseEstimator, TransformerMixin):\n    def fit(self, X, y=None):\n        return self\n\n    def transform(self, X):\n        return np.array([self.extract_vggish_embeddings(file_path) for file_path in X])\n    \n    def extract_vggish_embeddings(self, file_path):\n        sample, _ = librosa.load(file_path, sr=16000)\n        sample, _ = librosa.effects.trim(sample, top_db=10)\n        if len(sample) < 5 * 16000:\n            sample = np.pad(sample, (0, 5 * 16000 - len(sample)))\n        else:\n            sample = sample[:5 * 16000]\n        if np.max(np.abs(sample)) > 0:\n            sample = sample / np.max(np.abs(sample))\n        embeddings = vggish_model(sample)\n        embeddings_flattened = embeddings.numpy().flatten()\n        return embeddings_flattened\n\n# HuBERT feature extraction\nclass HubertFeatureExtractor(BaseEstimator, TransformerMixin):\n    def fit(self, X, y=None):\n        return self\n\n    def transform(self, X):\n        return np.array([self.extract_features(path) for path in X])\n\n    def extract_features(self, file_path):\n        sample, sr = librosa.load(file_path, sr=16000)\n        sample, _ = librosa.effects.trim(sample, top_db=10)\n        \n        if len(sample) < 5 * 16000:\n            sample = np.pad(sample, (0, 5 * 16000 - len(sample)))\n        else:\n            sample = sample[:5 * 16000]\n\n        if np.max(np.abs(sample)) > 0:\n            sample = sample / np.max(np.abs(sample))\n\n        input_values = processor(sample, sampling_rate=16000, return_tensors=\"pt\").input_values\n        with torch.no_grad():\n            hidden_states = hubert_model(input_values).last_hidden_state\n        features = hidden_states.mean(dim=1).squeeze().numpy()\n\n        return features\n\n# MFCC feature extraction\nclass MFCCFeatureExtractor(BaseEstimator, TransformerMixin):\n    def __init__(self, n_mfcc=40, sr=16000):\n        self.n_mfcc = n_mfcc\n        self.sr = sr\n\n    def fit(self, X, y=None):\n        return self\n\n    def transform(self, X):\n        return np.array([self.extract_features(path) for path in X])\n\n    def extract_features(self, file_path):\n        y, sr = librosa.load(file_path, sr=self.sr)\n        y, _ = librosa.effects.trim(y, top_db=10)\n\n        if len(y) < 5 * sr:\n            y = np.pad(y, (0, 5 * sr - len(y)))\n        else:\n            y = y[:5 * sr]\n\n        if np.max(np.abs(y)) > 0:\n            y = y / np.max(np.abs(y))\n\n        mfccs = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=self.n_mfcc)\n        return mfccs.mean(axis=1)\n\n# LogMel feature extraction\nclass LogMelFeatureExtractor(BaseEstimator, TransformerMixin):\n    def __init__(self, n_mels=64, sr=16000):\n        self.n_mels = n_mels\n        self.sr = sr\n\n    def fit(self, X, y=None):\n        return self\n\n    def transform(self, X):\n        return np.array([self.extract_features(path) for path in X])\n\n    def extract_features(self, file_path):\n        y, sr = librosa.load(file_path, sr=self.sr)\n        y, _ = librosa.effects.trim(y, top_db=10)\n\n        if len(y) < 5 * sr:\n            y = np.pad(y, (0, 5 * sr - len(y)))\n        else:\n            y = y[:5 * sr]\n\n        if np.max(np.abs(y)) > 0:\n            y = y / np.max(np.abs(y))\n\n        mel_spectrogram = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=self.n_mels)\n        log_mel_spectrogram = librosa.power_to_db(mel_spectrogram)\n        return log_mel_spectrogram.mean(axis=1)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Pipeline","metadata":{}},{"cell_type":"code","source":"# Create the pipeline for to apply feature extraction\n# Extracted features were flattened and will be converted into a combined numpy array to create a final X_train/test feature set\n# This pipleine will also be used to transfrom the test data for final submission\nfrom sklearn.pipeline import FeatureUnion, Pipeline\n\ncombined_features = FeatureUnion([\n    (\"vggish\", VGGishFeatureExtractor()),\n    (\"mfcc\", MFCCFeatureExtractor(n_mfcc=40)),\n    (\"hubert\", HubertFeatureExtractor())\n    #(\"logmel\", LogMelFeatureExtractor(n_mels=64)) ## commented out since best results do not include these features\n])\n\nfeature_pipeline = Pipeline([\n    (\"features\", combined_features)\n])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Create Training Feature Arrays\n- Here we apply the feature extraction pipeline on the training files to create the extracted features for training","metadata":{}},{"cell_type":"code","source":"train_audio_paths = [os.path.join(audio_dir, fname) for fname in train_files['fname']]\ntest_audio_paths = [os.path.join(audio_dir, fname) for fname in test_files['fname']]\n\n#Combine all extracted features into train and test sets\nX_train_combined = feature_pipeline.fit_transform(train_audio_paths)\nX_test_combined = feature_pipeline.transform(test_audio_paths)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Prepare Features for Training","metadata":{}},{"cell_type":"code","source":"X_train_combined.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_test_combined.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_files.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Save Features, Pipeline","metadata":{}},{"cell_type":"code","source":"#Save X_train/X_test for use in training\nnp.save('X_train_vmh_mv_final.npy', X_train_combined)\nnp.save('X_test_vmh_mv_final.npy', X_test_combined)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save original dataframes for creating y_tarain/y_test in the training steps\ntrain_files.to_pickle('train_mv_final.pkl')\ntest_files.to_pickle('test_mv_final.pkl')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save feature extraction pipleine to use for final evaluation prep\njoblib.dump(feature_pipeline, 'feature_extraction_pipeline_mv_vmh_final.joblib')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}