{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":13955246,"sourceType":"datasetVersion","datasetId":8895159}],"dockerImageVersionId":31192,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\"\"\"\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename).head())\"\"\"\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:18.195620Z","iopub.execute_input":"2025-12-07T13:33:18.196669Z","iopub.status.idle":"2025-12-07T13:33:20.439841Z","shell.execute_reply.started":"2025-12-07T13:33:18.196598Z","shell.execute_reply":"2025-12-07T13:33:20.438727Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport glob\nimport numpy as np\nimport pandas as pd\n\nimport librosa\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import accuracy_score, classification_report, confusion_matrix\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom sklearn.metrics import accuracy_score, classification_report, confusion_matrix\nfrom tensorflow.keras import models","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:20.441476Z","iopub.execute_input":"2025-12-07T13:33:20.442115Z","iopub.status.idle":"2025-12-07T13:33:45.061015Z","shell.execute_reply.started":"2025-12-07T13:33:20.442083Z","shell.execute_reply":"2025-12-07T13:33:45.059719Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\ndef plot_confusion_matrix(y_true, y_pred, labels, title):\n    cm = confusion_matrix(y_true, y_pred, labels=list(range(len(labels))))\n    plt.figure(figsize=(8, 6))\n\n    sns.heatmap(\n        cm,\n        annot=True,\n        fmt=\"g\",\n        cmap=\"magma\",\n        linewidths=.5,\n        cbar=True,\n        xticklabels=labels,\n        yticklabels=labels\n    )\n\n    plt.xlabel(\"Predicted\", fontsize=12)\n    plt.ylabel(\"True\", fontsize=12)\n    plt.title(title, fontsize=14)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:45.062076Z","iopub.execute_input":"2025-12-07T13:33:45.063040Z","iopub.status.idle":"2025-12-07T13:33:45.688772Z","shell.execute_reply.started":"2025-12-07T13:33:45.063010Z","shell.execute_reply":"2025-12-07T13:33:45.687701Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# INTRODUCTION ","metadata":{}},{"cell_type":"markdown","source":"This project focuses on building a **speech-command recognition system** capable of identifying a set of **twelve spoken words** from the TensorFlow Speech Recognition Challenge dataset. Since the scope of our course is **introductory machine learning**, we first approached the problem using *classical ML models* such as *Logistic Regression* and *Random Forest*. These baseline methods allowed us to understand the limitations of traditional algorithms when dealing with raw audio features.\n\nTo extend the analysis and explore more powerful architectures, we then transitioned to **Deep Learning models**. Specifically, we implemented a *Multilayer Perceptron* **(MLP)** and a *Convolutional Neural Network* **(CNN)**, the latter being naturally well-suited for spectrogram-based audio classification. Our *objective* was to **compare all four models**, evaluate their strengths and weaknesses, and determine how performance scales as we move from simple linear methods to more complex neural networks.\n\nThis step-by-step strategy enabled us to ***highlight the importance of feature representation and model capacity*** in speech recognition tasks, while keeping the project consistent with the progression of the course.","metadata":{}},{"cell_type":"markdown","source":"# 1. Data preparation and MFCC feature extraction (classical models)","metadata":{}},{"cell_type":"markdown","source":"## 1.1 Dataset paths and label set\n","metadata":{}},{"cell_type":"code","source":"DATA_DIR = \"/kaggle/input/train0/train/audio\" \n\n\nCORE_COMMANDS = [\n    \"yes\", \"no\",\n    \"up\", \"down\",\n    \"left\", \"right\",\n    \"on\", \"off\",\n    \"stop\", \"go\"\n]\n\n# we consider 2 special classes in our labels \nSPECIAL_LABELS = [\"unknown\", \"silence\"]\n\n\nALL_LABELS = CORE_COMMANDS + SPECIAL_LABELS\n\nprint(\"Classes used:\", ALL_LABELS)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:45.690786Z","iopub.execute_input":"2025-12-07T13:33:45.691434Z","iopub.status.idle":"2025-12-07T13:33:45.961369Z","shell.execute_reply.started":"2025-12-07T13:33:45.691405Z","shell.execute_reply":"2025-12-07T13:33:45.959778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#We extracted all .wav files in the DATA_DIR and count them\nfilepaths = glob.glob(os.path.join(DATA_DIR, \"*\", \"*.wav\"))\nlen(filepaths)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:45.962575Z","iopub.execute_input":"2025-12-07T13:33:45.962968Z","iopub.status.idle":"2025-12-07T13:33:47.896769Z","shell.execute_reply.started":"2025-12-07T13:33:45.962930Z","shell.execute_reply":"2025-12-07T13:33:47.895642Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.2 Build a dataframe of audio files","metadata":{}},{"cell_type":"code","source":"# build a table with path ,label as described and fname (file name)\n\nrows = []\n\nfor path in filepaths:\n    \n    parts = path.split(os.sep)\n    folder = parts[-2]  # accordind to the pattern name , we just take the part related to the label for each files\n    \n    if folder in CORE_COMMANDS:\n        label = folder\n    elif folder == \"_background_noise_\":\n        label = \"silence\"\n    else:\n        label = \"unknown\"\n    \n    rows.append((os.path.basename(path), label, path))\n\ndf = pd.DataFrame(rows, columns=[\"fname\", \"label\", \"path\"])\n\n\n\n#checking the aspect of the new created dataframe\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:47.898106Z","iopub.execute_input":"2025-12-07T13:33:47.898467Z","iopub.status.idle":"2025-12-07T13:33:48.043452Z","shell.execute_reply.started":"2025-12-07T13:33:47.898434Z","shell.execute_reply":"2025-12-07T13:33:48.042356Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.3 MFCC feature extraction function","metadata":{}},{"cell_type":"markdown","source":"According to our ML project, we use an **MFCC-based feature** extraction approach.\nThe sample rate is set to **16 kHz** because the Speech Commands dataset is recorded at this frequency.\nThis means that for one second of audio, approximately **16,000 samples** are captured.\n\n**MFCC (Mel-Frequency Cepstral Coefficients)** is a representation of speech that compresses the spectral information of the signal in a way that matches human auditory perception.\nInstead of working directly with raw audio, the signal is split into short frames, and for each frame, n_mfcc coefficients are computed.\nSince we set n_mfcc = 20, each frame is represented by 20 MFCC values.\nFinally, we summarize the entire audio clip by taking the mean and standard deviation over all frames, resulting in a fixed-size feature vector of 40 values.","metadata":{}},{"cell_type":"code","source":"# Function to extract MFCC\n\nSAMPLE_RATE = 16000\nN_MFCC = 20              # number of MFCC coefficient\n\ndef extract_features(path, sr=SAMPLE_RATE, n_mfcc=N_MFCC):\n    # charge audio file\n    y, sr = librosa.load(path, sr=sr)\n    \n    # MFCC (shape: (n_mfcc, n_frames))\n    mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=n_mfcc)\n    \n    # we take average (mean) and standar-deviation (std) -> 2 * n_mfcc features\n    mfcc_mean = mfcc.mean(axis=1)\n    mfcc_std  = mfcc.std(axis=1)\n    \n    features = np.concatenate([mfcc_mean, mfcc_std], axis=0)  # shape: (2 * n_mfcc,)\n    \n    return features\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:48.044515Z","iopub.execute_input":"2025-12-07T13:33:48.045315Z","iopub.status.idle":"2025-12-07T13:33:48.052758Z","shell.execute_reply.started":"2025-12-07T13:33:48.045286Z","shell.execute_reply":"2025-12-07T13:33:48.051311Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ! Silence files Management","metadata":{}},{"cell_type":"code","source":"def extract_features_from_chunk(y_chunk, sr=SAMPLE_RATE, n_mfcc=N_MFCC):\n    mfcc = librosa.feature.mfcc(y=y_chunk, sr=sr, n_mfcc=n_mfcc)\n    mfcc_mean = mfcc.mean(axis=1)\n    mfcc_std = mfcc.std(axis=1)\n    features = np.concatenate([mfcc_mean, mfcc_std], axis=0)  # shape (40,)\n    return features","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:48.053855Z","iopub.execute_input":"2025-12-07T13:33:48.054137Z","iopub.status.idle":"2025-12-07T13:33:48.081735Z","shell.execute_reply.started":"2025-12-07T13:33:48.054112Z","shell.execute_reply":"2025-12-07T13:33:48.080251Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def extract_silence_chunks(path, \n                           sr=SAMPLE_RATE, \n                           n_mfcc=N_MFCC,\n                           chunk_duration=1.0,   \n                           hop_duration=0.5):    \n    \n    y, sr = librosa.load(path, sr=sr)\n\n    chunk_len = int(chunk_duration * sr)\n    hop_len   = int(hop_duration * sr)\n\n    features_list = []\n\n    if len(y) < chunk_len:\n        \n        features_list.append(extract_features_from_chunk(y, sr=sr, n_mfcc=n_mfcc))\n        return features_list\n\n    for start in range(0, len(y) - chunk_len + 1, hop_len):\n        y_chunk = y[start:start + chunk_len]\n        feats = extract_features_from_chunk(y_chunk, sr=sr, n_mfcc=n_mfcc)\n        features_list.append(feats)\n\n    return features_list","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:48.082797Z","iopub.execute_input":"2025-12-07T13:33:48.083128Z","iopub.status.idle":"2025-12-07T13:33:48.108340Z","shell.execute_reply.started":"2025-12-07T13:33:48.083107Z","shell.execute_reply":"2025-12-07T13:33:48.107015Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.4 Build feature matrix X and label vector y","metadata":{}},{"cell_type":"code","source":"from tqdm.auto import tqdm   \ndf_sample = df.copy()\n\nX_list = []\ny_list = []\n\nlabel_to_idx = {label: i for i, label in enumerate(ALL_LABELS)}\nidx_to_label = {i: label for label, i in label_to_idx.items()}\n\n\nfor _, row in tqdm(df_sample.iterrows(),\n                   total=len(df_sample),\n                   desc=\"Extraction des features MFCC\"):\n    path = row[\"path\"]\n    label_str = row[\"label\"]\n\n    if label_str not in ALL_LABELS:\n        continue\n\n    if label_str == \"silence\":\n        silence_feats_list = extract_silence_chunks(path)\n        for feats in silence_feats_list:\n            X_list.append(feats)\n            y_list.append(label_to_idx[\"silence\"])\n    else:\n        features = extract_features(path)\n        X_list.append(features)\n        y_list.append(label_to_idx[label_str])\n\nX = np.vstack(X_list)   # shape: (n_samples, n_features)\ny = np.array(y_list)\n\nX.shape, y.shape\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:33:48.112966Z","iopub.execute_input":"2025-12-07T13:33:48.113411Z","iopub.status.idle":"2025-12-07T13:55:34.688948Z","shell.execute_reply.started":"2025-12-07T13:33:48.113377Z","shell.execute_reply":"2025-12-07T13:55:34.687200Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For the classical ML models, we used **MFCC features compressed** by taking the **mean** and **standard deviation** across time.\nThis representation allows models such as **Logistic Regression, Random Forest or MLP** to process audio, since they require **fixed-size vectors**.\nHowever, this compression completely removes the **temporal structure** of speech, which is essential for distinguishing phonemes and short commands.\nAs a consequence, classical models cannot capture the dynamics of speech and get **limited performance**","metadata":{}},{"cell_type":"code","source":"\"\"\"\n#Code to build a subset for classsical models (LogReg, RF, MLP) to limit preprocessing time\n#we limit the number of exemples per class \n\n\nMAX_PER_CLASS = 1000   # UP to 1800 for better trained model\ndf_limited = (\n    df.groupby(\"label\", group_keys=False)\n      .apply(lambda x: x.sample(min(len(x), MAX_PER_CLASS), random_state=0))\n      .reset_index(drop=True)\n)\n\nprint(\"Nb d'exemples après limitation :\", len(df_limited))\ndf_limited[\"label\"].value_counts()\n\n\nX_list = []\ny_list = []\n\nlabel_to_idx = {label: i for i, label in enumerate(ALL_LABELS)}\nidx_to_label = {i: label for label, i in label_to_idx.items()}\n\nfrom tqdm import tqdm\n\nfor _, row in tqdm(df_limited.iterrows(), total=len(df_limited)):\n    path = row[\"path\"]\n    label_str = row[\"label\"]\n\n    if label_str not in ALL_LABELS:\n        continue\n\n    features = extract_features(path)   # MFCC mean/std\n    X_list.append(features)\n    y_list.append(label_to_idx[label_str])\n\nX = np.vstack(X_list)   # shape: (n_samples, n_features)\ny = np.array(y_list)\n\nX.shape, y.shape\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:34.690261Z","iopub.execute_input":"2025-12-07T13:55:34.691036Z","iopub.status.idle":"2025-12-07T13:55:34.701719Z","shell.execute_reply.started":"2025-12-07T13:55:34.691009Z","shell.execute_reply":"2025-12-07T13:55:34.700256Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.5 Train/validation split and standardization","metadata":{}},{"cell_type":"markdown","source":"***Why do we create our own train/validation split instead of using the Kaggle test set?***\n\nThe test set provided on ***Kaggle*** does ***not include labels***. Its purpose is solely to evaluate the final submission on the competition server.\nUsing it during training or model selection would therefore be impossible and, more importantly, methodologically incorrect.\nTo properly tune and compare our models, we **need a labeled subset** that the model has never seen during training.\nFor this reason, we ***split our available labeled data***.\nThis ensures that all performance metrics in this notebook are computed on data with known labels that remain fully independent from training.","metadata":{}},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(\n    X, y,\n    test_size=0.2,\n    random_state=0,\n    stratify=y\n)\n\nscaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_val_scaled   = scaler.transform(X_val)\n\nX_train_scaled.shape, X_val_scaled.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:34.703083Z","iopub.execute_input":"2025-12-07T13:55:34.703517Z","iopub.status.idle":"2025-12-07T13:55:34.899239Z","shell.execute_reply.started":"2025-12-07T13:55:34.703486Z","shell.execute_reply":"2025-12-07T13:55:34.898211Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. LOGISTIC REGRESSION","metadata":{}},{"cell_type":"markdown","source":"Logistic Regression is a simple linear classifier that serves as a baseline model for multiclass audio classification.\nIt assumes that each class can be separated by linear decision boundaries in the feature space.\nDespite its simplicity, it is useful to evaluate whether the extracted MFCC features contain enough discriminative information for a linear model to perform well.\nWe therefore train a multinomial Logistic Regression on the MFCC features and evaluate its performance on the validation set.\n\nAlthough MFCCs are originally designed for speech processing, reducing them to simple statistical summaries such as the mean and standard deviation creates a compact, low-dimensional representation of each audio file. This transformation removes most of the temporal variability and keeps only the global spectral signature of the word.\n\nIn this context, a linear classifier like Logistic Regression becomes a reasonable baseline:\n\neach sample is represented by only a few dozen features (e.g., 40 values),\n\nthese features describe the global energy distribution across frequency bands,\n\ndifferent spoken commands may exhibit distinct average spectral profiles,\n\nand linear decision boundaries can sometimes be sufficient to separate such compact representations.\n\nTherefore, MFCC(mean+std) provides a simplified feature space where Logistic Regression could perform adequately if the global spectral characteristics of the commands were linearly separable.","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Model implementation (LogReg)","metadata":{}},{"cell_type":"code","source":"logreg = LogisticRegression(\n    max_iter=8000,\n    multi_class=\"multinomial\",\n    n_jobs=-1\n)\n\n################################       Training Time      ####################################################\nimport time\nstart = time.time()\n\nlogreg.fit(X_train_scaled, y_train)\n\nend = time.time()\nprint(f\"Training time: {end - start:.2f} seconds\")\n#############################################################################################################","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:34.900148Z","iopub.execute_input":"2025-12-07T13:55:34.900443Z","iopub.status.idle":"2025-12-07T13:55:53.439034Z","shell.execute_reply.started":"2025-12-07T13:55:34.900422Z","shell.execute_reply":"2025-12-07T13:55:53.437393Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.2 Evaluation","metadata":{}},{"cell_type":"code","source":"y_val_pred_lr = logreg.predict(X_val_scaled)\n\nacc_lr = accuracy_score(y_val, y_val_pred_lr)\nprint(\"Logistic Regression - Validation Accuracy:\", acc_lr)\nprint(classification_report(y_val, y_val_pred_lr, target_names=ALL_LABELS))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:53.440366Z","iopub.execute_input":"2025-12-07T13:55:53.440635Z","iopub.status.idle":"2025-12-07T13:55:53.479744Z","shell.execute_reply.started":"2025-12-07T13:55:53.440616Z","shell.execute_reply":"2025-12-07T13:55:53.478752Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.3 Vizualisation","metadata":{}},{"cell_type":"markdown","source":"### PCA 2D ","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\nimport matplotlib.pyplot as plt\nimport numpy as np\n\n\npca = PCA(n_components=2, random_state=0)\nX_val_2d = pca.fit_transform(X_val_scaled)\n\n\nclasses_to_show = [\"up\", \"down\", \"on\", \"off\", \"unknown\"]\n\nplt.figure(figsize=(8, 6))\nfor label in classes_to_show:\n    idx = np.where(y_val == ALL_LABELS.index(label))[0]\n    plt.scatter(\n        X_val_2d[idx, 0],\n        X_val_2d[idx, 1],\n        s=10,\n        alpha=0.5,\n        label=label\n    )\n\nplt.title(\"PCA of MFCC(mean/std) features (validation set)\")\nplt.xlabel(\"PC1\")\nplt.ylabel(\"PC2\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:53.480689Z","iopub.execute_input":"2025-12-07T13:55:53.481027Z","iopub.status.idle":"2025-12-07T13:55:54.285841Z","shell.execute_reply.started":"2025-12-07T13:55:53.480998Z","shell.execute_reply":"2025-12-07T13:55:54.283514Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Probability Courbs LogReg","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n\nprobas_val = logreg.predict_proba(X_val_scaled)  # shape (N, 12)\n\n\nwords = [\"up\", \"down\", \"off\", \"unknown\"]\nindices = [ALL_LABELS.index(w) for w in words]\n\n\ntrue_up_idx = np.where(y_val == ALL_LABELS.index(\"up\"))[0][:50]  # 50 premiers\n\nplt.figure(figsize=(8, 4))\nfor j, (w, idx_cls) in enumerate(zip(words, indices)):\n    plt.plot(\n        probas_val[true_up_idx, idx_cls],\n        label=f\"P({w} | x)\",\n        marker=\"o\",\n        linestyle=\"-\",\n        alpha=0.7\n    )\n\nplt.xlabel(\"Sample index (true class = 'up')\")\nplt.ylabel(\"Predicted probability\")\nplt.title(\"LogReg probabilities for different words (true 'up')\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:54.287585Z","iopub.execute_input":"2025-12-07T13:55:54.288398Z","iopub.status.idle":"2025-12-07T13:55:54.600477Z","shell.execute_reply.started":"2025-12-07T13:55:54.288355Z","shell.execute_reply":"2025-12-07T13:55:54.597310Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.4 LogReg Confusion Matrix","metadata":{}},{"cell_type":"code","source":"plot_confusion_matrix(\n    y_val, \n    y_val_pred_lr, \n    ALL_LABELS,\n    \"Confusion Matrix – Logistic Regression\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:54.603251Z","iopub.execute_input":"2025-12-07T13:55:54.603598Z","iopub.status.idle":"2025-12-07T13:55:55.214131Z","shell.execute_reply.started":"2025-12-07T13:55:54.603573Z","shell.execute_reply":"2025-12-07T13:55:55.212988Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Results & Conclusion\n\nDespite its simplicity, this approach suffers from an important limitation:\nby reducing each MFCC sequence to only two statistics (mean and standard deviation), all temporal information of the audio signal is removed.\n\nSpeech commands, even very short ones, are highly dynamic signals where the evolution of the spectrum over time carries essential information. Two different words may share similar average frequency content but differ in how the energy moves across frequencies during articulation. Once this temporal structure is flattened into global averages, these differences become nearly invisible to a linear model.\n\nAs a consequence:\n\nmany commands become linearly indistinguishable,\n\nsubtle phonetic patterns are lost,\n\nthe model struggles especially on short or noisy words,\n\nand the classifier tends to collapse on majority classes such as unknown.\n\nThis explains why Logistic Regression cannot fully exploit the richness of the audio data and achieves limited performance, despite the MFCC representation containing relevant spectral information.","metadata":{}},{"cell_type":"markdown","source":"### How could logistic regression be improved?","metadata":{}},{"cell_type":"markdown","source":"Logistic Regression could be improved mainly by providing it with richer audio features. Instead of using only MFCC means and standard deviations which remove all temporal information we could include:\nMFCC matrices, delta and delta-delta coefficients because it give more significant differences of courbs between classs representation.\nHowever, even with these improvements, Logistic Regression remains limited by its linear nature and cannot fully model the temporal dynamics of speech.","metadata":{}},{"cell_type":"markdown","source":"# 3. RANDOM FOREST","metadata":{}},{"cell_type":"markdown","source":"Random Forest is an **ensemble learning method** based on the aggregation of multiple decision **trees trained** on different bootstrap samples of the data. Each individual tree is a weak, high-variance classifier, but averaging hundreds of them reduces variance and improves robustness.\n\nThis model is particularly relevant for evaluating how well **non-linear methods** can exploit MFCC features. Unlike Logistic Regression, which can only learn linear boundaries, Random Forests can model **more complex**, non-linear interactions between MFCC coefficients. With MFCC(mean/std), the input space contains global spectral patterns that may not be linearly separable, and decision trees are capable of performing threshold-based splits that capture such relationships. This makes Random Forests **theoretically** better suited than Logistic Regression to exploit these global acoustic differences.\n\nIn addition, Random Forests are robust to noise, require minimal feature preprocessing, and often perform well on medium-sized tabular datasets—making them a reasonable candidate for MFCC-based audio classification.\n\nWe train a Random Forest classifier with **300** trees and balanced class weights, then evaluate it on the validation set.","metadata":{}},{"cell_type":"markdown","source":"## 3.1 Model implementation (RFC)","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nrf = RandomForestClassifier(\n    n_estimators=300,         \n    max_depth=None,          \n    min_samples_split=2,\n    min_samples_leaf=1,\n    max_features=\"sqrt\",    \n    n_jobs=-1,\n    random_state=0,\n    class_weight=\"balanced\"\n)\n\n################################       Training Time      ####################################################\nimport time\nstart = time.time()\n\nrf.fit(X_train_scaled, y_train)\n\nend = time.time()\nprint(f\"Training time: {end - start:.2f} seconds\")\n#############################################################################################################","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:55:55.215282Z","iopub.execute_input":"2025-12-07T13:55:55.215561Z","iopub.status.idle":"2025-12-07T13:57:00.874570Z","shell.execute_reply.started":"2025-12-07T13:55:55.215539Z","shell.execute_reply":"2025-12-07T13:57:00.873469Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.2 Evaluation","metadata":{}},{"cell_type":"code","source":"y_val_pred_rf = rf.predict(X_val)\n\nacc_rf = accuracy_score(y_val, y_val_pred_rf)\nprint(\"Random Forest - Validation Accuracy:\", acc_rf)\nprint(classification_report(y_val, y_val_pred_rf, target_names=ALL_LABELS))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:57:00.875749Z","iopub.execute_input":"2025-12-07T13:57:00.876155Z","iopub.status.idle":"2025-12-07T13:57:01.241832Z","shell.execute_reply.started":"2025-12-07T13:57:00.876130Z","shell.execute_reply":"2025-12-07T13:57:01.240811Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.3 Vizualisation","metadata":{}},{"cell_type":"markdown","source":"### Probability Courbs","metadata":{}},{"cell_type":"code","source":"words = [\"up\", \"down\", \"off\", \"unknown\"]\nindices = [ALL_LABELS.index(w) for w in words]\n\ntrue_up_idx = np.where(y_val == ALL_LABELS.index(\"up\"))[0][:50]\n\nprobas_val_rf = rf.predict_proba(X_val_scaled)\n\nplt.figure(figsize=(8, 4))\nfor w, idx_cls in zip(words, indices):\n    plt.plot(\n        probas_val_rf[true_up_idx, idx_cls],\n        marker=\"o\", linestyle=\"-\", alpha=0.7,\n        label=f\"P({w} | x) RF\"\n    )\n\nplt.xlabel(\"Sample index (true class = 'up')\")\nplt.ylabel(\"Predicted probability\")\nplt.title(\"Random Forest probabilities for different words (true 'up')\")\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:57:01.243219Z","iopub.execute_input":"2025-12-07T13:57:01.243550Z","iopub.status.idle":"2025-12-07T13:57:02.141802Z","shell.execute_reply.started":"2025-12-07T13:57:01.243528Z","shell.execute_reply":"2025-12-07T13:57:02.140622Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Features Importance","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\nimportances = rf.feature_importances_\nindices = np.argsort(importances)[::-1]\n\nplt.figure(figsize=(8, 4))\nplt.bar(range(10), importances[indices[:10]])\nplt.xticks(range(10), [f\"f{idx}\" for idx in indices[:10]], rotation=45)\nplt.xlabel(\"Top features\")\nplt.ylabel(\"Importance\")\nplt.title(\"Random Forest feature importances\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:57:02.142876Z","iopub.execute_input":"2025-12-07T13:57:02.143221Z","iopub.status.idle":"2025-12-07T13:57:02.443228Z","shell.execute_reply.started":"2025-12-07T13:57:02.143198Z","shell.execute_reply":"2025-12-07T13:57:02.442144Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.3 RF Confusion Matrix","metadata":{}},{"cell_type":"code","source":"plot_confusion_matrix(\n    y_val, \n    y_val_pred_rf, \n    ALL_LABELS,\n    \"Confusion Matrix – Random Forest\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:57:02.444313Z","iopub.execute_input":"2025-12-07T13:57:02.444684Z","iopub.status.idle":"2025-12-07T13:57:02.986475Z","shell.execute_reply.started":"2025-12-07T13:57:02.444630Z","shell.execute_reply":"2025-12-07T13:57:02.985435Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Conclusion\n\nEven though Random Forests are powerful non-linear models, they collapse completely on this audio classification task.\nThe MFCC(mean/std) representation removes all temporal structure and compresses each audio clip into ~80 global statistics. For speech commands, this representation is too weak and too homogeneous for tree-based models to learn meaningful splits.\n\nAs a result:\n\nall commands share very similar global MFCC statistics,\n\nthe Random Forest cannot find stable, informative splits in the feature space,\n\nThe trees default to the majority class (“unknown”), because it dominates the dataset,\n\nthe confusion matrix shows no correct predictions for any real command,\n\nand the overall accuracy is an illusion driven almost entirely by the dominance of unknown samples.\n\nIn short, Random Forests need richer and more structured features to be effective. With MFCC(mean/std), the information available is too limited and too homogeneous, causing the model to fail despite its non-linear capacity.","metadata":{}},{"cell_type":"markdown","source":"### How could Random Forest be improved?\n\nRandom Forest performance could be improved mainly by providing more informative audio features. Instead of using only MFCC means and standard deviations which are too compressed and discard temporal structure.                                                                      \nwe could extract richer representations such as full **MFCC sequences**, **delta** and **delta-delta coefficients**.These features would offer the model more discriminative spectral information and allow it to learn more meaningful splits.\n\nAnother improvement would be to ***reduce the dominance of the unknown class***, either by undersampling it or balancing the dataset more carefully. Since Random Forest tends to favor majority classes when features are not very informative, controlling the size of unknown would help prevent the model from collapsing into predicting that class for most samples.\n\nEven with these improvements, Random Forest remains limited by its tabular nature, but its performance could be made noticeably more stable and less biased than in the baseline version.","metadata":{}},{"cell_type":"markdown","source":"# 4. MultiLayer Perceptron (MLP)","metadata":{}},{"cell_type":"markdown","source":"A ***Multilayer Perceptron (MLP)*** is **one of the *simplest* neural network architectures**.\nIt operates on flattened feature vectors and performs a sequence of linear transformations followed by non-linear activations.\n\nIn this project, we use MFCC statistical features (mean and standard deviation), which produce a compact numerical vector for each audio file.\nThis representation is well suited for an MLP because it removes the temporal structure and allows the model to focus on global frequency patterns.\n\nCompared to Logistic Regression, the MLP can learn non-linear decision boundaries, making it a stronger classifier for tabular audio features.\nHowever, unlike CNNs, it does not exploit the 2-D spectrogram structure.\n\nOur MLP consists of several dense layers with ReLU activations, batch normalization and dropout to improve stability and reduce overfitting.\nWe also use early stopping to keep the best model during training.","metadata":{}},{"cell_type":"markdown","source":"## 4.1 labels and input dimensions","metadata":{}},{"cell_type":"code","source":"num_classes = len(ALL_LABELS)\ninput_dim = X_train_scaled.shape[1]\n\ny_train_cat = keras.utils.to_categorical(y_train, num_classes=num_classes)\ny_val_cat   = keras.utils.to_categorical(y_val,   num_classes=num_classes)\n\ny_train_cat.shape, y_val_cat.shape\ninput_dim","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:57:02.987447Z","iopub.execute_input":"2025-12-07T13:57:02.987723Z","iopub.status.idle":"2025-12-07T13:57:03.002236Z","shell.execute_reply.started":"2025-12-07T13:57:02.987702Z","shell.execute_reply":"2025-12-07T13:57:03.001090Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Since the ***MLP*** uses a softmax output layer combined with the ***categorical_crossentropy loss***, the class labels must be converted from integer form to one-hot encoded vectors.\nThis representation is required for multiclass neural networks because each **output neuron** corresponds to **one class**, and the **loss function** ***compares probability distributions*** rather than *integer class indices*.","metadata":{}},{"cell_type":"markdown","source":"## 4.2 Model implementation (MLP)","metadata":{}},{"cell_type":"code","source":"def build_mlp(input_dim, num_classes):\n    model = models.Sequential([\n        layers.Input(shape=(input_dim,)),\n        layers.Dense(256, activation='relu'),\n        layers.BatchNormalization(),\n        layers.Dropout(0.3),\n\n        layers.Dense(128, activation='relu'),\n        layers.BatchNormalization(),\n        layers.Dropout(0.3),\n\n        layers.Dense(64, activation='relu'),\n        layers.Dropout(0.2),\n\n        layers.Dense(num_classes, activation='softmax')\n    ])\n\n    model.compile(\n        optimizer='adam',\n        loss='categorical_crossentropy',\n        metrics=['accuracy']\n    )\n    return model\n\n#Model building\nmlp_model = build_mlp(input_dim, num_classes)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:57:03.003626Z","iopub.execute_input":"2025-12-07T13:57:03.003903Z","iopub.status.idle":"2025-12-07T13:57:03.168742Z","shell.execute_reply.started":"2025-12-07T13:57:03.003884Z","shell.execute_reply":"2025-12-07T13:57:03.167477Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.3 Training the MLP model","metadata":{}},{"cell_type":"code","source":"es = EarlyStopping(\n    monitor=\"val_accuracy\",\n    patience=3,\n    restore_best_weights=True\n)\n\n\nimport time\nstart = time.time()\n\n\nhistory = mlp_model.fit(\n    X_train_scaled, y_train_cat,\n    validation_data=(X_val_scaled, y_val_cat),\n    epochs=25,\n    batch_size=64,\n    callbacks=[es],\n    verbose=1\n)\n\n\n\nend = time.time()\nprint(f\"Training time: {end - start:.2f} seconds\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:57:03.169637Z","iopub.execute_input":"2025-12-07T13:57:03.169976Z","iopub.status.idle":"2025-12-07T13:58:31.353080Z","shell.execute_reply.started":"2025-12-07T13:57:03.169953Z","shell.execute_reply":"2025-12-07T13:58:31.352043Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.4 Training curves (Accuracy & Loss)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,4))\n\nplt.subplot(1,2,1)\nplt.plot(history.history[\"accuracy\"], label=\"train\")\nplt.plot(history.history[\"val_accuracy\"], label=\"val\")\nplt.title(\"MLP Accuracy\")\nplt.legend()\n\nplt.subplot(1,2,2)\nplt.plot(history.history[\"loss\"], label=\"train\")\nplt.plot(history.history[\"val_loss\"], label=\"val\")\nplt.title(\"MLP Loss\")\nplt.legend()\n\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:58:31.354194Z","iopub.execute_input":"2025-12-07T13:58:31.354559Z","iopub.status.idle":"2025-12-07T13:58:31.748628Z","shell.execute_reply.started":"2025-12-07T13:58:31.354524Z","shell.execute_reply":"2025-12-07T13:58:31.747684Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.5 Evaluation","metadata":{}},{"cell_type":"code","source":"y_val_pred_mlp = mlp_model.predict(X_val_scaled).argmax(axis=1)\n\nacc_mlp = accuracy_score(y_val, y_val_pred_mlp)\nprint(\"MLP - Validation Accuracy:\", acc_mlp)\nprint(classification_report(y_val, y_val_pred_mlp, target_names=ALL_LABELS))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:58:31.749577Z","iopub.execute_input":"2025-12-07T13:58:31.749943Z","iopub.status.idle":"2025-12-07T13:58:32.957738Z","shell.execute_reply.started":"2025-12-07T13:58:31.749920Z","shell.execute_reply":"2025-12-07T13:58:32.956467Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.6 Confusion Matrix","metadata":{}},{"cell_type":"code","source":"plot_confusion_matrix(\n    y_val,\n    y_val_pred_mlp,\n    ALL_LABELS,\n    \"Confusion Matrix – MLP\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:58:32.962532Z","iopub.execute_input":"2025-12-07T13:58:32.962875Z","iopub.status.idle":"2025-12-07T13:58:33.521677Z","shell.execute_reply.started":"2025-12-07T13:58:32.962849Z","shell.execute_reply":"2025-12-07T13:58:33.520293Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Conclusion\n\nThe MLP clearly performs better than Logistic Regression and Random Forest on our MFCC(mean/std) features. Thanks to its non-linear layers, it can capture relationships between MFCC coefficients that linear or tree-based models fail to learn. This leads to a validation accuracy of roughly 72–73%, with stable training curves and limited overfitting.\n\nHowever, the model remains limited by the MFCC(mean/std) representation, which removes all temporal structure from the audio. Words with similar global spectral profiles (e.g., up, on, off, go) are still frequently confused, and the dominance of the unknown class also affects performance. As a result, while the MLP is stronger than classical models, it still cannot exploit the full richness of speech patterns.","metadata":{}},{"cell_type":"markdown","source":"# 5. Convultional Neuronous Network (CNN) on log-Mel Spectrograms","metadata":{}},{"cell_type":"markdown","source":"## 5.1 Log-Mel spectrogram parameters and feature extraction","metadata":{}},{"cell_type":"markdown","source":"We *convert each audio waveform* into a **2D log-Mel spectrogram with 40 Mel bands and a fixed temporal length of 64 frames.**\nShorter clips are padded with *zeros*, longer clips are *cropped*, so that the CNN always receives inputs of shape (40, 64, 1).","metadata":{}},{"cell_type":"code","source":"#Audio Constants\nSAMPLE_RATE = 16000\nN_MELS = 40\nHOP_LENGTH = 256\nN_FFT = 1024\nMAX_LEN = 64   #Number of time frames\n\n#Extraction Fonction\ndef extract_logmel_image(path,\n                         sr=SAMPLE_RATE,\n                         n_mels=N_MELS,\n                         n_fft=N_FFT,\n                         hop_length=HOP_LENGTH,\n                         max_len=MAX_LEN):\n    y, sr = librosa.load(path, sr=sr)\n\n    # melspectrogram (n_mels x n_frames)\n    melspec = librosa.feature.melspectrogram(\n        y=y, sr=sr,\n        n_mels=n_mels,\n        n_fft=n_fft,\n        hop_length=hop_length\n    )\n    logmel = librosa.power_to_db(melspec + 1e-6)\n\n    \n    if logmel.shape[1] < max_len:\n        pad_width = max_len - logmel.shape[1]\n        logmel = np.pad(logmel, ((0, 0), (0, pad_width)), mode=\"constant\")\n    elif logmel.shape[1] > max_len:\n        logmel = logmel[:, :max_len]\n\n    return logmel  # shape: (n_mels, max_len)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:58:33.522407Z","iopub.execute_input":"2025-12-07T13:58:33.522785Z","iopub.status.idle":"2025-12-07T13:58:33.531439Z","shell.execute_reply.started":"2025-12-07T13:58:33.522757Z","shell.execute_reply":"2025-12-07T13:58:33.530087Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#test to see image features & shapes\ntest_img = extract_logmel_image(df.iloc[0][\"path\"])\ntest_img.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:58:33.532639Z","iopub.execute_input":"2025-12-07T13:58:33.533049Z","iopub.status.idle":"2025-12-07T13:58:33.570191Z","shell.execute_reply.started":"2025-12-07T13:58:33.532966Z","shell.execute_reply":"2025-12-07T13:58:33.568719Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## *! SILENCE FILES MANAGEMENT","metadata":{}},{"cell_type":"markdown","source":"Due to the very low quantity of silent files , we decide to split original registering into small new file ton complete the silence set","metadata":{}},{"cell_type":"code","source":"def extract_silence_chunks_logmel(path,\n                                  chunk_duration=1.0,\n                                  hop_duration=0.5,\n                                  sr=SAMPLE_RATE,\n                                  n_mels=N_MELS,\n                                  hop_length=HOP_LENGTH,\n                                  n_fft=N_FFT,max_len=MAX_LEN):\n\n\n    #Cut Background noise files into segments                               \n    y, sr = librosa.load(path, sr=sr)\n    chunk_len = int(chunk_duration * sr)\n    hop_len   = int(hop_duration * sr)\n\n    chunks = []\n\n    \n    if len(y) < chunk_len:\n        logmel = extract_logmel_image(path)\n        return [logmel]\n\n  \n    for start in range(0, len(y) - chunk_len + 1, hop_len):\n        y_chunk = y[start:start+chunk_len]\n\n        \n        mel = librosa.feature.melspectrogram(\n            y=y_chunk,\n            sr=sr,\n            n_mels=n_mels,\n            hop_length=hop_length,\n            n_fft=n_fft\n        )\n        logmel = librosa.power_to_db(mel + 1e-6)\n\n        # pad/crop pour avoir (40, 64)\n        if logmel.shape[1] < max_len:\n            pad_width = max_len - logmel.shape[1]\n            logmel = np.pad(logmel, ((0, 0), (0, pad_width)), mode=\"constant\")\n        else:\n            logmel = logmel[:, :max_len]\n\n        chunks.append(logmel)\n\n    return chunks\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:58:33.571605Z","iopub.execute_input":"2025-12-07T13:58:33.572161Z","iopub.status.idle":"2025-12-07T13:58:33.582578Z","shell.execute_reply.started":"2025-12-07T13:58:33.572135Z","shell.execute_reply":"2025-12-07T13:58:33.581469Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.2 Building the CNN dataset (images + labels)","metadata":{}},{"cell_type":"markdown","source":"For **each audio file**, we compute its log-Mel spectrogram and stack all spectrograms into a 4D tensor X_img with ***shape (N, 40, 64, 1)***.          \n***The labels are mapped to integer indices in y_cnn.***","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nimport numpy as np\n\n\nlabel_to_idx = {label: i for i, label in enumerate(ALL_LABELS)}\nidx_to_label = {i: label for label, i in label_to_idx.items()}\n\n\ndf_silence = df[df[\"label\"] == \"silence\"]\ndf_not_silence = df[df[\"label\"] != \"silence\"]\n\n\nsilence_imgs = []\nsilence_labels = []\n\nfor path in df_silence[\"path\"]:\n    chunks = extract_silence_chunks_logmel(path)  # -> liste de (40, 64)\n    for img in chunks:\n        silence_imgs.append(img)\n        silence_labels.append(label_to_idx[\"silence\"])\n\nprint(\"Nb of log-mel silence :\", len(silence_imgs))\n\n\nX_img_list = []\ny_list = []\n\nfor _, row in df.iterrows():\n    label_str = row[\"label\"]\n    if label_str not in ALL_LABELS:\n        \n        continue\n\n    img = extract_logmel_image(row[\"path\"])  # (40, 64)\n    X_img_list.append(img)\n    y_list.append(label_to_idx[label_str])\n\n\nX_img_list.extend(silence_imgs)\ny_list.extend(silence_labels)\n\n\nX_img = np.stack(X_img_list).astype(\"float32\")   # (N, 40, 64)\ny_cnn = np.array(y_list, dtype=\"int64\")          # (N,)\n\n# 6) Ajouter la dimension canal pour Conv2D\nX_img = X_img[..., np.newaxis]                   # (N, 40, 64, 1)\n\nprint(\"Dataset global CNN :\", X_img.shape)\nunique_labels, counts = np.unique(y_cnn, return_counts=True)\nprint(\"global class distribution:\")\nfor lbl, c in zip(unique_labels, counts):\n    print(f\"{idx_to_label[lbl]:8s} -> {c}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T13:58:33.584131Z","iopub.execute_input":"2025-12-07T13:58:33.584496Z","iopub.status.idle":"2025-12-07T14:05:45.856064Z","shell.execute_reply.started":"2025-12-07T13:58:33.584472Z","shell.execute_reply":"2025-12-07T14:05:45.854109Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.3 Train/validation split","metadata":{}},{"cell_type":"markdown","source":"We keep a validation set to evaluate the CNN on unseen data.\n*As for the MLP*, the labels are converted to one-hot vectors because the model uses a **softmax output** layer with the **categorical_crossentropy loss**.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom tensorflow import keras\n\nX_train_img, X_val_img, y_train_cnn, y_val_cnn = train_test_split(\n    X_img, y_cnn,\n    test_size=0.2,\n    random_state=0,\n    stratify=y_cnn\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:05:45.857301Z","iopub.execute_input":"2025-12-07T14:05:45.857573Z","iopub.status.idle":"2025-12-07T14:05:46.122422Z","shell.execute_reply.started":"2025-12-07T14:05:45.857551Z","shell.execute_reply":"2025-12-07T14:05:46.120904Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sil_idx = label_to_idx[\"silence\"]\n\nprint(\"Silence in y_train_cnn :\", (y_train_cnn == sil_idx).sum())\nprint(\"Silence in y_val_cnn   :\", (y_val_cnn   == sil_idx).sum())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:05:46.123622Z","iopub.execute_input":"2025-12-07T14:05:46.124025Z","iopub.status.idle":"2025-12-07T14:05:46.132734Z","shell.execute_reply.started":"2025-12-07T14:05:46.123997Z","shell.execute_reply":"2025-12-07T14:05:46.131484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Input and labels dimensions\nnum_classes = len(ALL_LABELS)\n\ny_train_cat = keras.utils.to_categorical(y_train_cnn, num_classes=num_classes)\ny_val_cat   = keras.utils.to_categorical(y_val_cnn,   num_classes=num_classes)\n\nX_train_img.shape, X_val_img.shape, y_train_cat.shape, y_val_cat.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:05:46.133557Z","iopub.execute_input":"2025-12-07T14:05:46.133896Z","iopub.status.idle":"2025-12-07T14:05:46.158341Z","shell.execute_reply.started":"2025-12-07T14:05:46.133856Z","shell.execute_reply":"2025-12-07T14:05:46.157324Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.4 CNN Implementation","metadata":{}},{"cell_type":"markdown","source":"The CNN takes as **input a log-Mel spectrogram of shape (40, 64, 1)** and applies several convolution + max-pooling blocks, followed by a dense layer and a **softmax output layer**.\nConvolutions allow the model to exploit the 2D time–frequency structure of the spectrograms.","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras import layers, models\n\ninput_shape = X_train_img.shape[1:]  # (40, 64, 1)\n\ndef build_cnn(input_shape, num_classes):\n    model = models.Sequential([\n        layers.Conv2D(16, (3, 3), activation=\"relu\", padding=\"same\",\n                      input_shape=input_shape),\n        layers.MaxPooling2D((2, 2)),\n\n        layers.Conv2D(32, (3, 3), activation=\"relu\", padding=\"same\"),\n        layers.MaxPooling2D((2, 2)),\n\n        layers.Conv2D(64, (3, 3), activation=\"relu\", padding=\"same\"),\n        layers.MaxPooling2D((2, 2)),\n\n        layers.Flatten(),\n        layers.Dense(128, activation=\"relu\"),\n        layers.Dropout(0.3),\n        layers.Dense(num_classes, activation=\"softmax\")\n    ])\n\n    model.compile(\n        optimizer=keras.optimizers.Adam(learning_rate=1e-3),\n        loss=\"categorical_crossentropy\",\n        metrics=[\"accuracy\"]\n    )\n    return model\n\ncnn_model = build_cnn(input_shape, num_classes)\ncnn_model.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:05:46.159544Z","iopub.execute_input":"2025-12-07T14:05:46.160003Z","iopub.status.idle":"2025-12-07T14:05:46.275639Z","shell.execute_reply.started":"2025-12-07T14:05:46.159971Z","shell.execute_reply":"2025-12-07T14:05:46.274598Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.5 training the CNN Model","metadata":{}},{"cell_type":"code","source":"import time\nstart = time.time()\n\nhistory_cnn = cnn_model.fit(\n    X_train_img, y_train_cat,\n    validation_data=(X_val_img, y_val_cat),\n    epochs=20,           # up to 20-30 based on model convergence\n    batch_size=64,\n    verbose=1\n)\n\nend = time.time()\nprint(f\"Training time (CNN): {end - start:.2f} seconds\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:05:46.276723Z","iopub.execute_input":"2025-12-07T14:05:46.277072Z","iopub.status.idle":"2025-12-07T14:27:42.355955Z","shell.execute_reply.started":"2025-12-07T14:05:46.277043Z","shell.execute_reply":"2025-12-07T14:27:42.354518Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.6 Training curves (Accuracy & Loss)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12, 4))\n\nplt.subplot(1, 2, 1)\nplt.plot(history_cnn.history[\"accuracy\"], label=\"train\")\nplt.plot(history_cnn.history[\"val_accuracy\"], label=\"val\")\nplt.title(\"CNN Accuracy\")\nplt.legend()\n\nplt.subplot(1, 2, 2)\nplt.plot(history_cnn.history[\"loss\"], label=\"train\")\nplt.plot(history_cnn.history[\"val_loss\"], label=\"val\")\nplt.title(\"CNN Loss\")\nplt.legend()\n\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:42.357346Z","iopub.execute_input":"2025-12-07T14:27:42.357701Z","iopub.status.idle":"2025-12-07T14:27:42.758273Z","shell.execute_reply.started":"2025-12-07T14:27:42.357638Z","shell.execute_reply":"2025-12-07T14:27:42.757292Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.7 Evaluation of the model on the validation set ","metadata":{}},{"cell_type":"markdown","source":"We compute the validation accuracy and the classification report to compare the CNN with the previous models.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, classification_report, confusion_matrix\nimport numpy as np\n\ny_val_pred_cnn = cnn_model.predict(X_val_img).argmax(axis=1)\n\nacc_cnn = accuracy_score(y_val_cnn, y_val_pred_cnn)\nprint(\"Validation accuracy (CNN):\", acc_cnn)\nprint()\nprint(classification_report(\n    y_val_cnn,\n    y_val_pred_cnn,\n    labels=list(range(num_classes)),\n    target_names=ALL_LABELS,\n    zero_division=0\n))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:42.759303Z","iopub.execute_input":"2025-12-07T14:27:42.759542Z","iopub.status.idle":"2025-12-07T14:27:48.570714Z","shell.execute_reply.started":"2025-12-07T14:27:42.759525Z","shell.execute_reply":"2025-12-07T14:27:48.569416Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.8 Confusion Matrix","metadata":{}},{"cell_type":"markdown","source":"The confusion matrix highlights which commands are best recognized and which ones are more often confused by the CNN.","metadata":{}},{"cell_type":"code","source":"plot_confusion_matrix(\n    y_val_cnn,\n    y_val_pred_cnn,\n    ALL_LABELS,\n    \"Confusion Matrix – CNN\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:48.571886Z","iopub.execute_input":"2025-12-07T14:27:48.572860Z","iopub.status.idle":"2025-12-07T14:27:49.141760Z","shell.execute_reply.started":"2025-12-07T14:27:48.572832Z","shell.execute_reply":"2025-12-07T14:27:49.140531Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### How to improve?","metadata":{}},{"cell_type":"markdown","source":"We could improve model performance without increasing the quantity of layer or sample by using more performent spectogram : ","metadata":{}},{"cell_type":"markdown","source":"* **Mel-Spectrogram + Delta + Delta-Delta**\n* **CMVN (cepstral mean-variance normalization)**\n* **Frame-stacking**","metadata":{}},{"cell_type":"markdown","source":"# EVALUATION COMPARATION TABLE","metadata":{}},{"cell_type":"markdown","source":"The following table summarizes the validation accuracy obtained with all four models: Logistic Regression, Random Forest, MLP, and CNN.\nThis comparison highlights how performance improves as model complexity increases, especially when moving from classical ML methods to deep learning architectures.","metadata":{}},{"cell_type":"code","source":"results = [\n    (\"Logistic Regression\", acc_lr),\n    (\"Random Forest\", acc_rf),\n    (\"MLP\", acc_mlp),\n    (\"CNN\", acc_cnn),\n]\n\npd.DataFrame(results, columns=[\"Model\", \"Validation Accuracy\"])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:49.142787Z","iopub.execute_input":"2025-12-07T14:27:49.143126Z","iopub.status.idle":"2025-12-07T14:27:49.157956Z","shell.execute_reply.started":"2025-12-07T14:27:49.143091Z","shell.execute_reply":"2025-12-07T14:27:49.156784Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\ndf_results = pd.DataFrame(results, columns=[\"Model\", \"Validation Accuracy\"])\n\nplt.figure(figsize=(8,5))\nsns.barplot(data=df_results, x=\"Model\", y=\"Validation Accuracy\", palette=\"Blues_d\")\nplt.title(\"Model Performance Comparison\")\nplt.ylim(0,1)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:49.159001Z","iopub.execute_input":"2025-12-07T14:27:49.159335Z","iopub.status.idle":"2025-12-07T14:27:49.358347Z","shell.execute_reply.started":"2025-12-07T14:27:49.159311Z","shell.execute_reply":"2025-12-07T14:27:49.357139Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\nimport numpy as np\n\n# 1) Sauvegarder le modèle CNN\ncnn_model.save(\"cnn_speech_model.h5\")\n\n# 2) Sauvegarder la liste des labels dans l'ordre des indices\nnp.save(\"all_labels.npy\", np.array(ALL_LABELS))\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:49.359536Z","iopub.execute_input":"2025-12-07T14:27:49.359938Z","iopub.status.idle":"2025-12-07T14:27:49.367630Z","shell.execute_reply.started":"2025-12-07T14:27:49.359905Z","shell.execute_reply":"2025-12-07T14:27:49.366280Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# COMPARATIVE ANALYSIS: CNN WITH FULL UNKNOWN SET VS CNN WITH UNKNOWN CLASS REGULARISATION","metadata":{}},{"cell_type":"raw","source":"","metadata":{}},{"cell_type":"markdown","source":"In this section, we evaluate how our CNN behaves when the unknown class is heavily over-represented compared to when it is regularised and balanced with the other classes.\nThe goal is to assess whether reducing the dominance of the unknown category helps the model learn more discriminative boundaries and improves overall performance particularly for the actual speech command classes.","metadata":{}},{"cell_type":"code","source":"core_classes = [lbl for lbl in ALL_LABELS if lbl not in [\"unknown\", \"silence\"]]\n\n\ndf_core = df[df[\"label\"].isin(core_classes)]\n\n\ndf_unknown = df[df[\"label\"] == \"unknown\"]\n\n\ndf_silence = df[df[\"label\"] == \"silence\"]\n\n\nclass_counts_core = df_core[\"label\"].value_counts()\ntarget_unknown = class_counts_core.min()\n\nprint(\"Min among core speech classes :\", target_unknown)\n\n\ndf_unknown_sample = df_unknown.sample(\n    n=target_unknown,\n    random_state=42\n)\n\n# balanced df building\ndf_cnn_bal = pd.concat(\n    [df_core, df_silence, df_unknown_sample],  \n    ignore_index=True\n)\n\ndf_cnn_bal = df_cnn_bal.sample(frac=1, random_state=42).reset_index(drop=True)\n\nprint(\"Balanced df shape :\", df_cnn_bal.shape)\nprint(df_cnn_bal[\"label\"].value_counts())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:49.369142Z","iopub.execute_input":"2025-12-07T14:27:49.369460Z","iopub.status.idle":"2025-12-07T14:27:49.432791Z","shell.execute_reply.started":"2025-12-07T14:27:49.369438Z","shell.execute_reply":"2025-12-07T14:27:49.431495Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def build_cnn_dataset_from_df(df_source, silence_imgs, silence_labels, label_to_idx):\n    X_list = []\n    y_list = []\n\n    for _, row in df_source.iterrows():\n        label_str = row[\"label\"]\n        if label_str not in label_to_idx:\n            continue\n\n        img = extract_logmel_image(row[\"path\"])  # (40, 64)\n        X_list.append(img)\n        y_list.append(label_to_idx[label_str])\n\n    \n    X_list.extend(silence_imgs)\n    y_list.extend(silence_labels)\n\n    X = np.stack(X_list).astype(\"float32\")   # (N, 40, 64)\n    y = np.array(y_list, dtype=\"int64\")      # (N,)\n\n    \n    X = X[..., np.newaxis]                  # (N, 40, 64, 1)\n    return X, y\n\n\nX_img_bal, y_cnn_bal = build_cnn_dataset_from_df(\n    df_cnn_bal,\n    silence_imgs,\n    silence_labels,\n    label_to_idx\n)\n\nprint(\"BALANCED CNN dataset shape :\", X_img_bal.shape)\nu_b, c_b = np.unique(y_cnn_bal, return_counts=True)\nprint(\"BALANCED class distribution :\")\nfor lbl, cnt in zip(u_b, c_b):\n    print(f\"{idx_to_label[lbl]:8s} -> {cnt}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:27:49.434020Z","iopub.execute_input":"2025-12-07T14:27:49.434384Z","iopub.status.idle":"2025-12-07T14:30:35.976283Z","shell.execute_reply.started":"2025-12-07T14:27:49.434361Z","shell.execute_reply":"2025-12-07T14:30:35.975288Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom tensorflow.keras.utils import to_categorical\n\nXtr_bal, Xval_bal, ytr_bal, yval_bal = train_test_split(\n    X_img_bal,\n    y_cnn_bal,\n    test_size=0.2,\n    random_state=42,\n    stratify=y_cnn_bal\n)\n\nprint(\"Train shape (bal) :\", Xtr_bal.shape)\nprint(\"Val shape (bal)   :\", Xval_bal.shape)\n\nnum_classes = len(ALL_LABELS)\n\nytr_bal_cat = to_categorical(ytr_bal, num_classes=num_classes)\nyval_bal_cat = to_categorical(yval_bal, num_classes=num_classes)\n\nprint(\"ytr_bal_cat shape :\", ytr_bal_cat.shape)\nprint(\"yval_bal_cat shape:\", yval_bal_cat.shape)\n\n\n\ncnn_bal = build_cnn(\n    input_shape=(40, 64, 1),\n    num_classes=len(ALL_LABELS)\n)\n\nhistory_bal = cnn_bal.fit(\n    Xtr_bal, ytr_bal_cat,\n    validation_data=(Xval_bal, yval_bal_cat),\n    epochs=20,\n    batch_size=64,\n    verbose=1  \n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:38:20.214574Z","iopub.execute_input":"2025-12-07T14:38:20.216387Z","iopub.status.idle":"2025-12-07T14:47:11.425317Z","shell.execute_reply.started":"2025-12-07T14:38:20.216346Z","shell.execute_reply":"2025-12-07T14:47:11.424288Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import classification_report, confusion_matrix, accuracy_score\n\n\nyval_pred_bal = np.argmax(cnn_bal.predict(Xval_bal), axis=1)\n\n\nacc_bal = accuracy_score(yval_bal, yval_pred_bal)\nprint(f\"\\nOverall Accuracy – CNN unknown reduced : {acc_bal:.4f}\\n\")\n\n\nprint(\"=== Classification report – CNN unknown reduced ===\")\nprint(classification_report(yval_bal, yval_pred_bal, target_names=ALL_LABELS))\n\n\nplot_confusion_matrix(\n    yval_bal,\n    yval_pred_bal,\n    ALL_LABELS,\n    \"Confusion Matrix – CNN (unknown reduced)\"\n)\n\n############################################################\nprint(\"=== Classification report – CNN full unknown ===\")\nprint(\"Validation accuracy (CNN):\", acc_cnn)\nprint()\nprint(classification_report(\n    y_val_cnn,\n    y_val_pred_cnn,\n    labels=list(range(num_classes)),\n    target_names=ALL_LABELS,\n    zero_division=0\n))\n\nplot_confusion_matrix(\n    y_val_cnn,\n    y_val_pred_cnn,\n    ALL_LABELS,\n    \"Confusion Matrix – CNN (full unknown)\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T15:47:59.470687Z","iopub.execute_input":"2025-12-07T15:47:59.471167Z","iopub.status.idle":"2025-12-07T15:48:02.976945Z","shell.execute_reply.started":"2025-12-07T15:47:59.471140Z","shell.execute_reply":"2025-12-07T15:48:02.975800Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Comparison Between Full-Unknown CNN and Reduced-Unknown CNN","metadata":{}},{"cell_type":"markdown","source":"### 1. Data Distribution","metadata":{}},{"cell_type":"markdown","source":"Full-unknown CNN\n* Contains 8208 “unknown” samples.\n* Very large variety of noises / out-of-vocabulary words.\n* All other classes (~475 samples each) become strong minorities → highly unbalanced dataset.\n                                                                                          \nReduced-unknown CNN\n* Contains ≈470 “unknown” samples.\n* Dataset is balanced across all classes.\n                                       \nConclusion:\nThe full-unknown dataset is strongly imbalanced, while the reduced version is balanced.\nThis directly affects how both models behave during training.\n\n","metadata":{}},{"cell_type":"markdown","source":"### 2. Global Performance","metadata":{}},{"cell_type":"markdown","source":"* Full-unknown CNN accuracy: ~95.6%\n* Reduced-unknown CNN accuracy: ~92.7%\n\nGlobal accuracy is not a reliable indicator here, because it is heavily influenced by the “unknown” class, which is extremely frequent in the full-unknown dataset.","metadata":{}},{"cell_type":"markdown","source":"### 3. Per-Class Analysis","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T15:11:35.899988Z","iopub.execute_input":"2025-12-07T15:11:35.900784Z","iopub.status.idle":"2025-12-07T15:11:35.991393Z","shell.execute_reply.started":"2025-12-07T15:11:35.900748Z","shell.execute_reply":"2025-12-07T15:11:35.989177Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# RANDOM FOREST WITH BALANCED UNKNOWN CLASS","metadata":{}},{"cell_type":"markdown","source":"Since RF is a model suffering by the majority class ; we decide to test this model on a balanced dataset to see if it can give better performance than the first ; the first was only able to recognise unknown and used to put any worlld in this class.\nWe use the same random forest model , but we just use a balanced dataset with unkonw class ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n\ndf_feats = pd.DataFrame(X)\ndf_feats[\"label\"] = y\n\n\nlabel_to_idx = {label: i for i, label in enumerate(ALL_LABELS)}\nidx_to_label = {i: label for label, i in label_to_idx.items()}\n\nunknown_idx = label_to_idx[\"unknown\"]\nsilence_idx = label_to_idx[\"silence\"]\n\ncore_idxs = [label_to_idx[lbl] for lbl in ALL_LABELS if lbl not in [\"unknown\", \"silence\"]]\n\n\ndf_core = df_feats[df_feats[\"label\"].isin(core_idxs)]\ndf_unknown = df_feats[df_feats[\"label\"] == unknown_idx]\ndf_silence = df_feats[df_feats[\"label\"] == silence_idx]\n\n\nmin_count = df_core[\"label\"].value_counts().min()\nprint(\"Min count among core speech classes :\", min_count)\n\ndf_unknown_sample = df_unknown.sample(n=min_count, random_state=0)\n\n\ndf_rf_bal = pd.concat(\n    [df_core, df_silence, df_unknown_sample],\n    ignore_index=True\n)\ndf_rf_bal = df_rf_bal.sample(frac=1, random_state=0).reset_index(drop=True)\n\nprint(\"Balanced label counts (RF):\")\nprint(df_rf_bal[\"label\"].map(idx_to_label).value_counts())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:55:05.853195Z","iopub.execute_input":"2025-12-07T14:55:05.853634Z","iopub.status.idle":"2025-12-07T14:55:05.927358Z","shell.execute_reply.started":"2025-12-07T14:55:05.853607Z","shell.execute_reply":"2025-12-07T14:55:05.925819Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\n\n# Features / labels pour RF balanced\nX_bal = df_rf_bal.drop(columns=[\"label\"]).values\ny_bal = df_rf_bal[\"label\"].values\n\nX_train_bal, X_val_bal, y_train_bal, y_val_bal = train_test_split(\n    X_bal, y_bal,\n    test_size=0.2,\n    random_state=0,\n    stratify=y_bal\n)\n\nscaler_bal = StandardScaler()\nX_train_bal_scaled = scaler_bal.fit_transform(X_train_bal)\nX_val_bal_scaled   = scaler_bal.transform(X_val_bal)\n\nX_train_bal_scaled.shape, X_val_bal_scaled.shape\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:55:13.462733Z","iopub.execute_input":"2025-12-07T14:55:13.463161Z","iopub.status.idle":"2025-12-07T14:55:13.535621Z","shell.execute_reply.started":"2025-12-07T14:55:13.463135Z","shell.execute_reply":"2025-12-07T14:55:13.534669Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nimport time\n\nrf_bal = RandomForestClassifier(\n    n_estimators=300,\n    max_depth=None,\n    min_samples_split=2,\n    min_samples_leaf=1,\n    max_features=\"sqrt\",\n    n_jobs=-1,\n    random_state=0,\n    class_weight=\"balanced\"   # on garde le même setting\n)\n\nprint(\"==== Training RF with reduced unknown class ====\")\nstart = time.time()\nrf_bal.fit(X_train_bal_scaled, y_train_bal)\nend = time.time()\nprint(f\"Training time (RF unknown reduced): {end - start:.2f} seconds\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T14:55:17.513486Z","iopub.execute_input":"2025-12-07T14:55:17.514278Z","iopub.status.idle":"2025-12-07T14:55:41.302874Z","shell.execute_reply.started":"2025-12-07T14:55:17.514245Z","shell.execute_reply":"2025-12-07T14:55:41.301748Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, classification_report, confusion_matrix\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\ny_val_pred_rf_bal = rf_bal.predict(X_val_bal_scaled)\n\nacc_rf_bal = accuracy_score(y_val_bal, y_val_pred_rf_bal)\nprint(\"Random Forest (unknown reduced) - Validation Accuracy:\", acc_rf_bal)\n\nprint(\"\\n=== Classification Report – Random Forest (unknown reduced) ===\")\nprint(classification_report(\n    y_val_bal,\n    y_val_pred_rf_bal,\n    target_names=ALL_LABELS\n))\n\ncm_rf_bal = confusion_matrix(y_val_bal, y_val_pred_rf_bal)\n\nplt.figure(figsize=(8, 7))\nsns.heatmap(\n    cm_rf_bal,\n    annot=True,\n    fmt=\"d\",\n    cmap=\"magma\",\n    xticklabels=ALL_LABELS,\n    yticklabels=ALL_LABELS\n)\nplt.xlabel(\"Predicted\")\nplt.ylabel(\"True\")\nplt.title(\"Confusion Matrix – Random Forest (unknown reduced)\")\nplt.show()\n\n#########################################################################\nprint(\"\\n=== Classification Report – Random Forest (unknown full) ===\")\nprint(\"Random Forest - Validation Accuracy:\", acc_rf)\nprint(classification_report(y_val, y_val_pred_rf, target_names=ALL_LABELS))\n\nplot_confusion_matrix(\n    y_val, \n    y_val_pred_rf, \n    ALL_LABELS,\n    \"Confusion Matrix – Random Forest\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-07T15:54:00.428418Z","iopub.execute_input":"2025-12-07T15:54:00.428821Z","iopub.status.idle":"2025-12-07T15:54:02.719964Z","shell.execute_reply.started":"2025-12-07T15:54:00.428794Z","shell.execute_reply":"2025-12-07T15:54:02.719028Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1. ***Dataset Balance Drives Everything*** <br>\n* Full-unknown RF sees >60% of samples labeled unknown → the model becomes completely **biased**.\n* Reduced-unknown RF has all **classes balanced** → the model learns real decision boundaries instead of collapsing into majority-class predictions.\n2. ***Full-Unknown RF Fails on All Command Classes***<br>\n* Full-unknown RF predicts almost everything as “unknown” → F1-scores for all commands = 0.00.\n* Only the unknown class is detected (≈0.77 F1), inflating global accuracy **artificially (≈62%)**.\n* This is a classic majority-class domination effect.\n3. ***Reduced-Unknown RF Performs Correctly but Modestly***<br>\n* F1-scores for commands range 0.53–0.72 → the model learns each class but stays far behind CNN performance.\n* Confusion matrix shows real mistakes, but at least the model makes **meaningful predictions**.\n* Balanced data removes **bias** but RF still struggles with MFCC(mean/std), which contain complex non-linear patterns better handled by neural networks.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}