{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ASR Case Study on TensorFlow Speech Recognition Challenge dataset\n\n**Introduction**\n\nAutomatic Speech Recognition (ASR) is a technology that enables computers to recognize and process human speech into text. In this case study, the TensorFlow Speech Recognition Challenge dataset is used to perform various speech processing tasks such as visualization, feature extraction, data augmentation, and evaluation to understand the fundamentals of speech recognition systems.\n\n**Objectives**\n\n* To understand the basic concepts of Automatic Speech Recognition (ASR).\n* To perform preprocessing and feature extraction on speech data.\n* To implement data augmentation and evaluation techniques on the TensorFlow Speech Recognition dataset.\n* To analyze the effect of different speech processing methods on audio signals.","metadata":{}},{"cell_type":"code","source":"!apt-get -qq install p7zip-full\n!7z x -y /kaggle/input/competitions/tensorflow-speech-recognition-challenge/train.7z -o/kaggle/working/","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:35:39.703742Z","iopub.execute_input":"2026-07-02T12:35:39.704468Z","iopub.status.idle":"2026-07-02T12:37:02.758325Z","shell.execute_reply.started":"2026-07-02T12:35:39.704417Z","shell.execute_reply":"2026-07-02T12:37:02.757100Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 1: Input Libraries & Audio Loading\n\n**Description**\n\n**What is the task?**\nLoad the required libraries and read an audio file from the TensorFlow Speech Commands dataset.\n\n**Conceptual Understanding**\nAudio processing requires libraries such as Librosa, NumPy, and Matplotlib to analyze and visualize speech signals.\n\n**Observation**\nThe dataset contains one-second .wav files sampled at 16 kHz and organized into folders based on spoken commands.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport librosa\nimport librosa.display\nimport os\n\nyes_folder = \"/kaggle/working/train/audio/yes\"\n\naudio_file = os.listdir(yes_folder)[0]\n\naudio_path = os.path.join(yes_folder, audio_file)\n\nprint(audio_path)\n\nsignal, sr = librosa.load(audio_path, sr=None)\n\nprint(\"Sampling Rate:\", sr)\nprint(\"Duration:\", len(signal)/sr)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:10:44.725404Z","iopub.execute_input":"2026-07-02T12:10:44.726353Z","iopub.status.idle":"2026-07-02T12:10:58.044023Z","shell.execute_reply.started":"2026-07-02T12:10:44.726319Z","shell.execute_reply":"2026-07-02T12:10:58.043215Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 2: Time-Domain Audio Visualization\n\n**Description**\n\n**What is the task?**\nVisualize the speech waveform and its amplitude variations.\n\n**Conceptual Understanding**\nA waveform represents how the amplitude changes over time.\n\n**Observation**\nThe speech signal shows varying amplitudes corresponding to the pronunciation of the spoken command.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,4))\nlibrosa.display.waveshow(signal, sr=sr)\nplt.title(\"Speech Waveform\")\nplt.xlabel(\"Time\")\nplt.ylabel(\"Amplitude\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:11:50.472544Z","iopub.execute_input":"2026-07-02T12:11:50.473419Z","iopub.status.idle":"2026-07-02T12:11:50.876333Z","shell.execute_reply.started":"2026-07-02T12:11:50.473386Z","shell.execute_reply":"2026-07-02T12:11:50.875659Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 3: Windowing & Framing\n\n**Description**\n\n**What is the task?**\nDivide the audio signal into small overlapping frames.\n\n**Conceptual Understanding**\nSpeech is non-stationary; framing makes it approximately stationary.\n\n**Observation**\nThe one-second clips are divided into several short frames.","metadata":{}},{"cell_type":"code","source":"frame_length = 512\nhop_length = 256\n\nframes = librosa.util.frame(\n    signal,\n    frame_length=frame_length,\n    hop_length=hop_length\n)\n\nprint(\"Frame Shape:\", frames.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:12:12.742881Z","iopub.execute_input":"2026-07-02T12:12:12.743566Z","iopub.status.idle":"2026-07-02T12:12:12.748534Z","shell.execute_reply.started":"2026-07-02T12:12:12.743535Z","shell.execute_reply":"2026-07-02T12:12:12.747904Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 4: Log-Mel Spectrogram\n\n**Description**\n\n**What is the task?**\nExtract Log-Mel Spectrogram features.\n\n**Conceptual Understanding**\nThe Log-Mel Spectrogram captures frequency information in a way similar to human hearing.\n\n**Observation**\nDifferent spoken commands exhibit different spectrogram patterns.","metadata":{}},{"cell_type":"code","source":"mel = librosa.feature.melspectrogram(\n    y=signal,\n    sr=sr\n)\n\nlog_mel = librosa.power_to_db(mel)\n\nplt.figure(figsize=(12,4))\nlibrosa.display.specshow(\n    log_mel,\n    sr=sr,\n    x_axis='time',\n    y_axis='mel'\n)\nplt.colorbar()\nplt.title(\"Log-Mel Spectrogram\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:12:39.948386Z","iopub.execute_input":"2026-07-02T12:12:39.949320Z","iopub.status.idle":"2026-07-02T12:12:41.224191Z","shell.execute_reply.started":"2026-07-02T12:12:39.949278Z","shell.execute_reply":"2026-07-02T12:12:41.223545Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 5: Time Shifting\n\n**Description**\n\n**What is the task?**\nShift the audio signal in time.\n\n**Conceptual Understanding**\nTime shifting increases dataset diversity.\n\n**Observation**\nThe word remains unchanged but starts at a different position.","metadata":{}},{"cell_type":"code","source":"shift = int(sr * 0.2)\n\nshifted_signal = np.roll(signal, shift)\n\nplt.figure(figsize=(12,4))\nlibrosa.display.waveshow(shifted_signal, sr=sr)\nplt.title(\"Time Shifted Signal\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:13:00.892384Z","iopub.execute_input":"2026-07-02T12:13:00.893329Z","iopub.status.idle":"2026-07-02T12:13:01.217340Z","shell.execute_reply.started":"2026-07-02T12:13:00.893297Z","shell.execute_reply":"2026-07-02T12:13:01.216724Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 6: Adding Noise\n\n**Description**\n\n**What is the task?**\nAdd random noise to the audio signal.\n\n**Conceptual Understanding**\nNoise augmentation improves robustness.\n\n**Observation**\nThe speech remains understandable despite added noise.","metadata":{}},{"cell_type":"code","source":"noise = np.random.randn(len(signal))\n\nnoisy_signal = signal + 0.005 * noise\n\nplt.figure(figsize=(12,4))\nlibrosa.display.waveshow(noisy_signal, sr=sr)\nplt.title(\"Noisy Signal\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:13:23.755293Z","iopub.execute_input":"2026-07-02T12:13:23.755727Z","iopub.status.idle":"2026-07-02T12:13:24.195644Z","shell.execute_reply.started":"2026-07-02T12:13:23.755676Z","shell.execute_reply":"2026-07-02T12:13:24.194635Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 7: Word Error Rate (WER)\n\n**Description**\n\n**What is the task?**\nEvaluate ASR performance.\n\n**Conceptual Understanding**\nWER measures recognition errors.\n\n**Observation**\nFor single-word commands, WER is usually 0 or 1.","metadata":{}},{"cell_type":"code","source":"!pip install jiwer\n\nfrom jiwer import wer\n\nreference = \"yes\"\nprediction = \"yas\"\n\nerror = wer(reference, prediction)\n\nprint(\"WER:\", error)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:13:47.180950Z","iopub.execute_input":"2026-07-02T12:13:47.181608Z","iopub.status.idle":"2026-07-02T12:13:53.225294Z","shell.execute_reply.started":"2026-07-02T12:13:47.181574Z","shell.execute_reply":"2026-07-02T12:13:53.224437Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 8: Audio Normalization\n\n**Description**\n\n**What is the task?**\nNormalize audio amplitude.\n\n**Conceptual Understanding**\nNormalization removes loudness variations.\n\n**Observation**\nAll recordings become more consistent.","metadata":{}},{"cell_type":"code","source":"normalized = librosa.util.normalize(signal)\n\nplt.figure(figsize=(12,4))\nlibrosa.display.waveshow(normalized, sr=sr)\nplt.title(\"Normalized Audio\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:14:20.903969Z","iopub.execute_input":"2026-07-02T12:14:20.904540Z","iopub.status.idle":"2026-07-02T12:14:21.242402Z","shell.execute_reply.started":"2026-07-02T12:14:20.904504Z","shell.execute_reply":"2026-07-02T12:14:21.241545Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 9: Silence Stripping\n\n**Description**\n\n**What is the task?**\nRemove silent portions.\n\n**Conceptual Understanding**\nSilence contains no useful speech information.\n\n**Observation**\nSome files contain leading and trailing silence.","metadata":{}},{"cell_type":"code","source":"trimmed_signal, index = librosa.effects.trim(signal)\n\nprint(\"Original Length:\", len(signal))\nprint(\"Trimmed Length:\", len(trimmed_signal))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:14:44.172028Z","iopub.execute_input":"2026-07-02T12:14:44.172799Z","iopub.status.idle":"2026-07-02T12:14:44.776585Z","shell.execute_reply.started":"2026-07-02T12:14:44.172757Z","shell.execute_reply":"2026-07-02T12:14:44.775770Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 10: Frequency Domain Visualization\n\n**Description**\n\n**What is the task?**\nAnalyze frequency components.\n\n**Conceptual Understanding**\nFFT converts signals from time to frequency domain.\n\n**Observation**\nSpeech energy is concentrated in lower frequencies.","metadata":{}},{"cell_type":"code","source":"fft = np.fft.fft(signal)\nfreq = np.fft.fftfreq(len(signal))\n\nplt.figure(figsize=(12,4))\nplt.plot(freq, np.abs(fft))\nplt.title(\"Frequency Domain\")\nplt.xlabel(\"Frequency\")\nplt.ylabel(\"Magnitude\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:15:10.009597Z","iopub.execute_input":"2026-07-02T12:15:10.010988Z","iopub.status.idle":"2026-07-02T12:15:10.136672Z","shell.execute_reply.started":"2026-07-02T12:15:10.010939Z","shell.execute_reply":"2026-07-02T12:15:10.135790Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 11: Text Label Preprocessing\n\n**Description**\n\n**What is the task?**\nConvert text labels into numbers.\n\n**Conceptual Understanding**\nMachine learning models require numeric labels.\n\n**Observation**\nEach folder corresponds to one speech command.","metadata":{}},{"cell_type":"code","source":"labels = os.listdir(\"/kaggle/working/train/audio\")\n\nfrom sklearn.preprocessing import LabelEncoder\n\nencoder = LabelEncoder()\n\nencoded_labels = encoder.fit_transform(labels)\n\nprint(\"Labels:\")\nprint(labels)\n\nprint(\"\\nEncoded Labels:\")\nprint(encoded_labels)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:16:06.810348Z","iopub.execute_input":"2026-07-02T12:16:06.811134Z","iopub.status.idle":"2026-07-02T12:16:06.817257Z","shell.execute_reply.started":"2026-07-02T12:16:06.811101Z","shell.execute_reply":"2026-07-02T12:16:06.816486Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 12: LPC Computation\n\n**Description**\n\n**What is the task?**\nExtract LPC coefficients.\n\n**Conceptual Understanding**\nLPC models the vocal tract characteristics.\n\n**Observation**\nDifferent words produce different LPC coefficients.","metadata":{}},{"cell_type":"code","source":"lpc_coeff = librosa.lpc(signal, order=16)\n\nprint(lpc_coeff)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:17:11.918032Z","iopub.execute_input":"2026-07-02T12:17:11.918688Z","iopub.status.idle":"2026-07-02T12:17:13.862916Z","shell.execute_reply.started":"2026-07-02T12:17:11.918656Z","shell.execute_reply":"2026-07-02T12:17:13.862047Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 13: Pitch Shifting\n\n**Description**\n\n**What is the task?**\nModify the pitch of the speech.\n\n**Conceptual Understanding**\nPitch shifting simulates different speakers.\n\n**Observation**\nThe word remains the same but sounds different.","metadata":{}},{"cell_type":"code","source":"pitched_signal = librosa.effects.pitch_shift(\n    signal,\n    sr=sr,\n    n_steps=4\n)\n\nplt.figure(figsize=(12,4))\nlibrosa.display.waveshow(pitched_signal, sr=sr)\nplt.title(\"Pitch Shifted Signal\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:17:38.340154Z","iopub.execute_input":"2026-07-02T12:17:38.340603Z","iopub.status.idle":"2026-07-02T12:17:39.166498Z","shell.execute_reply.started":"2026-07-02T12:17:38.340574Z","shell.execute_reply":"2026-07-02T12:17:39.165659Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 14: Speed Perturbation\n\n**Description**\n\n**What is the task?**\nChange the speaking rate.\n\n**Conceptual Understanding**\nDifferent people speak at different speeds.\n\n**Observation**\nThe command remains recognizable.","metadata":{}},{"cell_type":"code","source":"faster = librosa.effects.time_stretch(\n    signal,\n    rate=1.2\n)\n\nslower = librosa.effects.time_stretch(\n    signal,\n    rate=0.8\n)\n\nprint(\"Original:\", len(signal))\nprint(\"Faster:\", len(faster))\nprint(\"Slower:\", len(slower))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:17:59.326294Z","iopub.execute_input":"2026-07-02T12:17:59.327125Z","iopub.status.idle":"2026-07-02T12:17:59.345638Z","shell.execute_reply.started":"2026-07-02T12:17:59.327089Z","shell.execute_reply":"2026-07-02T12:17:59.345011Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 15: Audio Downsampling\n\n**Description**\n\n**What is the task?**\nReduce the sampling rate.\n\n**Conceptual Understanding**\nDownsampling decreases computational requirements.\n\n**Observation**\nSpeech quality remains acceptable at lower sampling rates.","metadata":{}},{"cell_type":"code","source":"downsampled = librosa.resample(\n    signal,\n    orig_sr=sr,\n    target_sr=8000\n)\n\nprint(\"Original Samples:\", len(signal))\nprint(\"Downsampled Samples:\", len(downsampled))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-07-02T12:18:20.862490Z","iopub.execute_input":"2026-07-02T12:18:20.862812Z","iopub.status.idle":"2026-07-02T12:18:20.868653Z","shell.execute_reply.started":"2026-07-02T12:18:20.862786Z","shell.execute_reply":"2026-07-02T12:18:20.867762Z"}},"outputs":[],"execution_count":null}]}