{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\n\nfor root, dirs, files in os.walk('/kaggle/input'):\n    print(root)\n    print(\"Files:\", len(files))\n    print(\"-\"*50)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:14.788998Z","iopub.execute_input":"2026-06-24T14:54:14.789387Z","iopub.status.idle":"2026-06-24T14:54:15.045389Z","shell.execute_reply.started":"2026-06-24T14:54:14.789358Z","shell.execute_reply":"2026-06-24T14:54:15.044272Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import os\n\nfor root, dirs, files in os.walk('/kaggle/input'):\n    for file in files[:10]:\n        print(os.path.join(root, file))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:15.047716Z","iopub.execute_input":"2026-06-24T14:54:15.048721Z","iopub.status.idle":"2026-06-24T14:54:15.061572Z","shell.execute_reply.started":"2026-06-24T14:54:15.048682Z","shell.execute_reply":"2026-06-24T14:54:15.060453Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Dataset Exploration and Audit\n\n## Objective\n\nBefore applying any speech processing techniques, I want to understand the\nstructure of the Bengali ASR dataset. This includes the number of audio files,\ntheir durations, sample rates, transcript lengths, and overall characteristics.\n\n## Why this is important\n\nUnderstanding dataset properties helps determine appropriate preprocessing\nstrategies and reveals potential challenges such as long recordings,\nimbalanced durations, silence segments, or transcript complexity.\n\n## Findings\n\nThe dataset contains 113 training audio files and 113 corresponding transcript files. The naming convention is consistent, with each audio file having a matching transcript file.\n\nThe dataset also contains 24 test audio files and a sample submission file. This structure confirms that the dataset is organized for supervised ASR training and evaluation.\n\nInitial inspection showed that the recordings are stored in WAV format and transcripts are stored as TXT files, making them straightforward to process using common speech-processing libraries.","metadata":{}},{"cell_type":"code","source":"import os\n\naudio_dir = \"/kaggle/input/competitions/dl-sprint-4-0-bengali-long-form-speech-recognition/transcription/transcription/train/audio\"\nannotation_dir = \"/kaggle/input/competitions/dl-sprint-4-0-bengali-long-form-speech-recognition/transcription/transcription/train/annotation\"\n\naudio_files = sorted([\n    f for f in os.listdir(audio_dir)\n    if f.endswith(\".wav\")\n])\n\nannotation_files = sorted([\n    f for f in os.listdir(annotation_dir)\n    if f.endswith(\".txt\")\n])\n\nprint(\"Training audio files:\", len(audio_files))\nprint(\"Annotation files:\", len(annotation_files))\n\nprint(\"\\nFirst 5 audio files:\")\nfor f in audio_files[:5]:\n    print(f)\n\nprint(\"\\nFirst 5 annotation files:\")\nfor f in annotation_files[:5]:\n    print(f)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-27T12:12:58.371147Z","iopub.execute_input":"2026-06-27T12:12:58.371492Z","iopub.status.idle":"2026-06-27T12:12:58.410457Z","shell.execute_reply.started":"2026-06-27T12:12:58.371461Z","shell.execute_reply":"2026-06-27T12:12:58.409206Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Audio Duration and Sample Rate Analysis\n\n### Objective\n\nTo measure the duration and sampling rate of all recordings.\n\n### Concept\n\nASR systems are sensitive to recording length and sampling frequency.\nUnderstanding the distribution of these properties helps determine\nwhether preprocessing such as downsampling or silence removal is needed.\n\n## Findings\n\nOut of 113 audio files, 112 were successfully analyzed and one file (train_089.wav) could not be processed.\n\nAll valid recordings use the same sampling rate of 16,000 Hz, indicating a standardized dataset format. The average recording duration is approximately 3491.86 seconds (58.2 minutes), confirming that this is a long-form speech recognition dataset.\n\nThe shortest recording is 2405.41 seconds, while the longest recording is 5268.34 seconds, showing substantial variation in recording length.","metadata":{}},{"cell_type":"code","source":"import soundfile as sf\nfrom tqdm import tqdm\n\nstats = []\nbad_files = []\n\nfor file in tqdm(audio_files):\n    path = os.path.join(audio_dir, file)\n\n    try:\n        info = sf.info(path)\n\n        stats.append({\n            \"file\": file,\n            \"duration_sec\": info.duration,\n            \"sample_rate\": info.samplerate\n        })\n\n    except Exception as e:\n        bad_files.append(file)\n\nprint(\"Bad files:\", bad_files)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-27T12:13:10.804524Z","iopub.execute_input":"2026-06-27T12:13:10.804844Z","iopub.status.idle":"2026-06-27T12:13:11.580734Z","shell.execute_reply.started":"2026-06-27T12:13:10.804816Z","shell.execute_reply":"2026-06-27T12:13:11.579767Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Audio Duration Distribution\n\n### Objective\n\nVisualize how audio durations vary across the dataset.\n\n## Findings\n\nThe duration distribution shows that most recordings are clustered around the dataset average of approximately 3500 seconds.\n\nThe median duration is 3514.42 seconds, which is very close to the mean duration, suggesting that recording lengths are relatively balanced and not heavily skewed.\n\nBecause most recordings are close to one hour in length, the dataset is expected to be more challenging than short-command or conversational speech datasets.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\n\ndf_audio = pd.DataFrame(stats)\n\nprint(df_audio.head())\nprint(df_audio.shape)\n\nplt.figure(figsize=(10,5))\n\nplt.hist(df_audio[\"duration_sec\"], bins=15)\n\nplt.title(\"Distribution of Audio Durations\")\nplt.xlabel(\"Duration (seconds)\")\nplt.ylabel(\"Count\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-27T12:13:30.727200Z","iopub.execute_input":"2026-06-27T12:13:30.727578Z","iopub.status.idle":"2026-06-27T12:13:31.367072Z","shell.execute_reply.started":"2026-06-27T12:13:30.727501Z","shell.execute_reply":"2026-06-27T12:13:31.365808Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Transcript Analysis\n\n### Objective\n\nTo understand transcript lengths and linguistic complexity.\n\n### Concept\n\nLonger transcripts generally indicate more challenging speech recognition\ntasks and may increase Word Error Rate.\n\n## Findings\n\nTranscript analysis revealed an average transcript length of 31,376 characters and 5,753 words per recording.\n\nThe longest transcript contains 54,579 characters and 9,838 words, while the shortest transcript contains 1,485 characters and 250 words.\n\nThese results indicate that the dataset contains long-form Bengali narrative speech with substantial linguistic content. The large transcript sizes suggest that ASR systems trained on this dataset must handle long contextual dependencies and extensive vocabularies.","metadata":{}},{"cell_type":"code","source":"transcript_stats = []\n\nfor txt_file in annotation_files:\n    \n    path = os.path.join(annotation_dir, txt_file)\n\n    with open(path, \"r\", encoding=\"utf-8\") as f:\n        text = f.read().strip()\n\n    transcript_stats.append({\n        \"file\": txt_file,\n        \"characters\": len(text),\n        \"words\": len(text.split())\n    })\n\ndf_text = pd.DataFrame(transcript_stats)\n\nprint(df_text.head())\n\nprint(\"\\nTranscript Summary\")\nprint(df_text.describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-27T12:14:05.731417Z","iopub.execute_input":"2026-06-27T12:14:05.731939Z","iopub.status.idle":"2026-06-27T12:14:06.495398Z","shell.execute_reply.started":"2026-06-27T12:14:05.731907Z","shell.execute_reply":"2026-06-27T12:14:06.494343Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Dataset Audit Findings\n\nAfter exploring the Bengali Long-Form Speech Recognition dataset, I found that the corpus contains 113 training recordings and 24 test recordings. Each training audio file is paired with a corresponding transcript file.\n\nThe recordings are significantly longer than traditional speech datasets. The average recording duration is approximately 3492 seconds (58.2 minutes), while the longest recording exceeds 87 minutes. This indicates that the dataset is composed of long-form speech such as lectures, interviews, discussions, or presentations.\n\nAll valid recordings use a sampling rate of 16,000 Hz, which is a standard sampling frequency for modern Automatic Speech Recognition systems. This consistency simplifies preprocessing because no initial resampling is required.\n\nTranscript analysis revealed an average transcript length of approximately 5753 words per recording. Some recordings contain nearly 10,000 words, making this a challenging long-context ASR task where transcription errors can accumulate over time.\n\nDuring data exploration, one corrupted audio file (train_089.wav) was detected and excluded from statistical analysis.\n","metadata":{}},{"cell_type":"code","source":"shortest_file = df_audio.loc[df_audio[\"duration_sec\"].idxmin()]\nlongest_file = df_audio.loc[df_audio[\"duration_sec\"].idxmax()]\n\nprint(\"Shortest:\")\nprint(shortest_file)\n\nprint(\"\\nLongest:\")\nprint(longest_file)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-27T12:14:17.349192Z","iopub.execute_input":"2026-06-27T12:14:17.349684Z","iopub.status.idle":"2026-06-27T12:14:17.358821Z","shell.execute_reply.started":"2026-06-27T12:14:17.349651Z","shell.execute_reply":"2026-06-27T12:14:17.357865Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"median_duration = df_audio[\"duration_sec\"].median()\n\ndf_audio[\"distance\"] = abs(df_audio[\"duration_sec\"] - median_duration)\n\nrepresentative = df_audio.sort_values(\"distance\").iloc[0]\n\nprint(representative)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-27T12:14:27.658540Z","iopub.execute_input":"2026-06-27T12:14:27.659377Z","iopub.status.idle":"2026-06-27T12:14:27.675260Z","shell.execute_reply.started":"2026-06-27T12:14:27.659343Z","shell.execute_reply":"2026-06-27T12:14:27.674138Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Bengali Long-Form Speech Recognition Dataset Analysis\n\n## Overview\n\nThis case study investigates a Bengali long-form Automatic Speech Recognition (ASR) dataset provided through the DL Sprint 4.0 competition. The objective is to explore speech preprocessing, feature extraction, visualization, and data augmentation techniques commonly used in ASR pipelines.\n\n## Dataset Summary\n\nThe dataset contains:\n\n- 113 training audio files\n- 113 transcription files\n- 24 test audio files\n- Audio sampling rate of 16 kHz\n- Long-form recordings averaging approximately 58 minutes in duration\n\n## Goals\n\nThe goals of this analysis are:\n\n1. Understand the acoustic properties of the dataset.\n2. Visualize speech signals in time and frequency domains.\n3. Apply common speech preprocessing techniques.\n4. Perform feature extraction for ASR applications.\n5. Evaluate augmentation methods used in speech recognition systems.","metadata":{}},{"cell_type":"markdown","source":"# Task 1: Input Libraries and Audio Loading\n\n## Objective\n\nThe objective of this task is to import the required Python libraries for audio signal processing and load a representative Bengali speech recording from the dataset. Audio loading is the first step in any Automatic Speech Recognition pipeline because all later preprocessing and feature extraction operations depend on the waveform representation.\n\n## Conceptual Understanding\n\nSpeech recordings are stored as digital signals consisting of amplitude samples collected at a specific sampling frequency. Loading the audio allows us to inspect its duration, sample rate, and waveform characteristics before applying transformations.\n\n## Findings from This Dataset\n\nBased on the dataset audit, all valid recordings use a sampling rate of 16,000 Hz. A representative recording (train_022.wav) was selected because its duration is close to the dataset median. This makes it suitable for demonstrating signal-processing techniques while remaining representative of the overall corpus.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nimport librosa\nimport librosa.display\n\nimport soundfile as sf\n\nfrom IPython.display import Audio\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Representative file selected from dataset audit\n\naudio_path = os.path.join(\n    audio_dir,\n    \"train_022.wav\"\n)\n\n# Load full recording\ny, sr = librosa.load(audio_path, sr=None)\n\nduration_sec = len(y) / sr\n\nprint(f\"File: train_022.wav\")\nprint(f\"Sample Rate: {sr} Hz\")\nprint(f\"Duration: {duration_sec:.2f} seconds\")\nprint(f\"Duration: {duration_sec/60:.2f} minutes\")\nprint(f\"Total Samples: {len(y):,}\")\n\n# Create a smaller analysis segment\nanalysis_duration = 30  # seconds\n\ny_segment = y[: analysis_duration * sr]\n\nprint(f\"\\nAnalysis Segment Length: {len(y_segment)/sr:.2f} sec\")\n\nAudio(y_segment, rate=sr)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:15.754450Z","iopub.execute_input":"2026-06-24T14:54:15.754899Z","iopub.status.idle":"2026-06-24T14:54:16.161005Z","shell.execute_reply.started":"2026-06-24T14:54:15.754858Z","shell.execute_reply":"2026-06-24T14:54:16.157450Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Task 1 Additional Inspection\n\n### Objective\n\nInspect the amplitude distribution of the loaded recording.","metadata":{}},{"cell_type":"code","source":"print(\"Minimum Amplitude:\", np.min(y_segment))\nprint(\"Maximum Amplitude:\", np.max(y_segment))\nprint(\"Mean Amplitude:\", np.mean(y_segment))\nprint(\"Standard Deviation:\", np.std(y_segment))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:16.162429Z","iopub.execute_input":"2026-06-24T14:54:16.162705Z","iopub.status.idle":"2026-06-24T14:54:16.171596Z","shell.execute_reply.started":"2026-06-24T14:54:16.162681Z","shell.execute_reply":"2026-06-24T14:54:16.170281Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Findings\n\nThe selected recording (train_022.wav) has a duration of approximately 58.84 minutes and contains more than 56 million audio samples. This confirms that the dataset consists of long-form speech recordings rather than short utterances.\n\nThe audio was recorded at 16,000 Hz, which is the standard sampling frequency used by many modern speech recognition systems.\n\nAmplitude analysis showed values ranging from approximately -0.72 to +0.72, indicating that the recording contains strong speech activity without excessive clipping. The mean amplitude is extremely close to zero, suggesting that the waveform is properly centered and does not contain significant DC offset. The standard deviation of approximately 0.102 indicates a healthy variation in speech energy throughout the recording.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 2: Time-Domain Audio Visualization (3D Plot)\n\n## Objective\n\nThe objective of this task is to visualize the speech waveform in the time domain using a three-dimensional representation. This helps identify variations in amplitude across consecutive frames and provides insight into speech activity patterns.\n\n## Conceptual Understanding\n\nSpeech signals are non-stationary, meaning their characteristics change over time. By visualizing framed audio segments in three dimensions, it becomes easier to observe amplitude variations, speech bursts, pauses, and temporal structure.\n","metadata":{}},{"cell_type":"code","source":"from mpl_toolkits.mplot3d import Axes3D\n\n# Use only first 5 seconds for visualization\nviz_audio = y_segment[:5 * sr]\n\nframe_size = 512\nhop_length = 256\n\nframes = librosa.util.frame(\n    viz_audio,\n    frame_length=frame_size,\n    hop_length=hop_length\n)\n\n# Limit frames for plotting speed\nframes = frames[:, :200]\n\nfig = plt.figure(figsize=(12, 7))\nax = fig.add_subplot(111, projection='3d')\n\nX = np.arange(frames.shape[1])\nY = np.arange(frames.shape[0])\n\nX, Y = np.meshgrid(X, Y)\n\nax.plot_surface(\n    X,\n    Y,\n    frames,\n    cmap=\"viridis\",\n    linewidth=0,\n    antialiased=False\n)\n\nax.set_title(\"3D Time-Domain Speech Visualization\")\nax.set_xlabel(\"Frame Number\")\nax.set_ylabel(\"Sample Index\")\nax.set_zlabel(\"Amplitude\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:16.172981Z","iopub.execute_input":"2026-06-24T14:54:16.173431Z","iopub.status.idle":"2026-06-24T14:54:17.312440Z","shell.execute_reply.started":"2026-06-24T14:54:16.173404Z","shell.execute_reply":"2026-06-24T14:54:17.311502Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Additional Waveform Visualization\n\nA traditional waveform plot is included to compare against the 3D representation and to identify speech and silence regions more clearly.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15,4))\n\nlibrosa.display.waveshow(\n    viz_audio,\n    sr=sr\n)\n\nplt.title(\"Speech Waveform (First 5 Seconds)\")\nplt.xlabel(\"Time (seconds)\")\nplt.ylabel(\"Amplitude\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:17.313637Z","iopub.execute_input":"2026-06-24T14:54:17.314725Z","iopub.status.idle":"2026-06-24T14:54:17.859176Z","shell.execute_reply.started":"2026-06-24T14:54:17.314519Z","shell.execute_reply":"2026-06-24T14:54:17.857876Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Findings\n\nThe 3D time-domain visualization revealed clear variations in speech energy throughout the analyzed audio segment. During the initial portion of the recording, the waveform exhibited relatively small amplitude peaks, indicating lower speech energy or quieter speech activity.\n\nAround the middle section of the visualization (approximately 60% of the plotted frames), significantly higher amplitude peaks became visible. This suggests a period of stronger speech activity, increased vocal intensity, or denser acoustic content. Between approximately 60% and 80% of the plotted region, the waveform maintained consistently high energy levels.\n\nTowards the end of the analyzed segment, the amplitude decreased slightly before rising again, indicating continued speech activity with varying intensity levels. The absence of long flat regions suggests that the selected segment contains mostly continuous speech with limited silence.\n\nThe traditional waveform plot confirmed the observations from the 3D visualization and demonstrated that speech energy is distributed unevenly over time, which is expected in natural long-form spoken language recordings.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 3: Windowing and Framing\n\n## Objective\n\nThe objective of this task is to divide the continuous speech signal into smaller overlapping frames and apply a window function. Speech signals change over time, but within a short interval they can be approximated as stationary, making analysis easier.\n\n## Conceptual Understanding\n\nASR systems typically process speech in short windows of approximately 20–30 milliseconds. Windowing reduces edge discontinuities and spectral leakage, improving the quality of feature extraction methods such as Mel spectrograms and LPC analysis.\n","metadata":{}},{"cell_type":"code","source":"# Create a processing segment for Tasks 3 and other tasks\n\nprocessing_audio = y[:10 * sr]\n\nprint(f\"Processing segment duration: {len(processing_audio)/sr:.2f} sec\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T15:01:24.552303Z","iopub.execute_input":"2026-06-24T15:01:24.552633Z","iopub.status.idle":"2026-06-24T15:01:24.559271Z","shell.execute_reply.started":"2026-06-24T15:01:24.552606Z","shell.execute_reply":"2026-06-24T15:01:24.557980Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"frame_length = int(0.025 * sr)   # 25 ms\nhop_length = int(0.010 * sr)     # 10 ms\n\nframes = librosa.util.frame(\n    processing_audio,\n    frame_length=frame_length,\n    hop_length=hop_length\n)\n\nwindow = np.hamming(frame_length)\n\nwindowed_frames = frames * window.reshape(-1, 1)\n\nprint(\"Frame Length:\", frame_length)\nprint(\"Hop Length:\", hop_length)\nprint(\"Number of Frames:\", frames.shape[1])\nprint(\"Samples per Frame:\", frames.shape[0])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:17.872787Z","iopub.execute_input":"2026-06-24T14:54:17.873353Z","iopub.status.idle":"2026-06-24T14:54:17.891036Z","shell.execute_reply.started":"2026-06-24T14:54:17.873312Z","shell.execute_reply":"2026-06-24T14:54:17.889787Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Graph","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,4))\n\nplt.plot(windowed_frames[:, 20])\n\nplt.title(\"Example Hamming Windowed Frame\")\nplt.xlabel(\"Sample Index\")\nplt.ylabel(\"Amplitude\")\n\nplt.grid(True)\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:17.892302Z","iopub.execute_input":"2026-06-24T14:54:17.892603Z","iopub.status.idle":"2026-06-24T14:54:18.089387Z","shell.execute_reply.started":"2026-06-24T14:54:17.892565Z","shell.execute_reply":"2026-06-24T14:54:18.087972Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Findings\n\nThe speech signal was divided into 998 overlapping frames using a frame length of 400 samples (25 milliseconds) and a hop length of 160 samples (10 milliseconds). These values are commonly used in Automatic Speech Recognition systems because speech can be assumed to be approximately stationary within short time intervals.\n\nAfter applying the Hamming window, the amplitude of each frame became largest near the center and gradually decreased toward both edges. This tapering effect reduces abrupt discontinuities at frame boundaries and minimizes spectral leakage during frequency-domain analysis.\n\nThe plotted frame exhibited oscillatory speech patterns with stronger amplitude in the middle region and weaker amplitude near the beginning and end of the frame. This confirms that the Hamming window was successfully applied and prepared the signal for subsequent feature extraction tasks.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 4: Feature Extraction Using Log-Mel Spectrogram\n\n## Objective\n\nThe objective of this task is to transform the speech waveform into a time-frequency representation using the Mel scale. The Log-Mel Spectrogram is widely used in Automatic Speech Recognition because it captures frequency information in a way that approximates human auditory perception.\n\n## Conceptual Understanding\n\nHuman hearing is more sensitive to lower frequencies than higher frequencies. The Mel scale compresses high-frequency information and preserves greater detail in lower frequencies. Applying a logarithmic transformation further compresses the dynamic range and highlights important speech patterns.","metadata":{}},{"cell_type":"code","source":"n_mels = 128\n\nmel_spec = librosa.feature.melspectrogram(\n    y=processing_audio,\n    sr=sr,\n    n_mels=n_mels,\n    n_fft=1024,\n    hop_length=512\n)\n\nlog_mel = librosa.power_to_db(\n    mel_spec,\n    ref=np.max\n)\n\nprint(\"Log-Mel Shape:\", log_mel.shape)\n\nplt.figure(figsize=(14,5))\n\nlibrosa.display.specshow(\n    log_mel,\n    sr=sr,\n    hop_length=512,\n    x_axis='time',\n    y_axis='mel'\n)\n\nplt.colorbar(format='%+2.0f dB')\nplt.title(\"Log-Mel Spectrogram\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:18.091092Z","iopub.execute_input":"2026-06-24T14:54:18.091623Z","iopub.status.idle":"2026-06-24T14:54:18.474247Z","shell.execute_reply.started":"2026-06-24T14:54:18.091590Z","shell.execute_reply":"2026-06-24T14:54:18.472947Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Minimum dB:\", np.min(log_mel))\nprint(\"Maximum dB:\", np.max(log_mel))\nprint(\"Mean dB:\", np.mean(log_mel))\nprint(\"Shape:\", log_mel.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:18.475461Z","iopub.execute_input":"2026-06-24T14:54:18.475806Z","iopub.status.idle":"2026-06-24T14:54:18.482966Z","shell.execute_reply.started":"2026-06-24T14:54:18.475771Z","shell.execute_reply":"2026-06-24T14:54:18.481907Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Findings\n\nThe generated Log-Mel Spectrogram has a shape of (128, 313), meaning that the speech signal was represented using 128 Mel-frequency bands across 313 time frames.\n\nThe spectrogram contained predominantly bright regions, indicating that the selected audio segment contains continuous speech activity with substantial acoustic energy. Most of the visible energy was concentrated in the lower and middle frequency regions, which is consistent with human speech production because vowels and many speech components primarily occupy lower frequencies.\n\nThe logarithmic energy values ranged from -80 dB to 0 dB, with an average value of approximately -55.38 dB. This wide dynamic range demonstrates the presence of both strong speech components and weaker acoustic regions.\n\nNear the end of the spectrogram, a noticeable dark vertical region appeared, indicating a temporary reduction in acoustic energy. This may correspond to a short pause, silence interval, or transition within the recording. Overall, the spectrogram confirms that the dataset contains rich speech information suitable for ASR feature extraction.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 5: Basic Data Augmentation Using Time Shifting\n\n## Objective\n\nThe objective of this task is to apply time-shifting augmentation to a speech signal. Time shifting moves the waveform forward or backward in time without changing the actual speech content. This technique is commonly used in speech recognition systems to improve robustness against variations in speech alignment.\n\n## Conceptual Understanding\n\nDuring real-world recordings, speech may begin at different positions within an audio file. Time-shifting augmentation simulates these variations by changing the starting location of the waveform while preserving the original information. This helps machine learning models become less sensitive to timing differences.","metadata":{}},{"cell_type":"code","source":"def time_shift(audio, shift_seconds=0.5):\n\n    shift_samples = int(sr * shift_seconds)\n\n    return np.roll(audio, shift_samples)\n\ny_shifted = time_shift(processing_audio)\n\nplt.figure(figsize=(12,4))\n\nplt.plot(processing_audio[:5000], label=\"Original\")\n\nplt.plot(y_shifted[:5000], label=\"Shifted\")\n\nplt.title(\"Original vs Time Shifted Signal\")\n\nplt.xlabel(\"Sample Index\")\n\nplt.ylabel(\"Amplitude\")\n\nplt.legend()\n\nplt.show()\n\nprint(\"Original Samples:\", len(processing_audio))\nprint(\"Shifted Samples:\", len(y_shifted))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:18.484418Z","iopub.execute_input":"2026-06-24T14:54:18.484831Z","iopub.status.idle":"2026-06-24T14:54:18.732783Z","shell.execute_reply.started":"2026-06-24T14:54:18.484792Z","shell.execute_reply":"2026-06-24T14:54:18.731739Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nThe time-shifted signal retained the same total number of samples (160,000 samples) as the original audio segment, confirming that no speech information was removed during augmentation.\n\nAfter applying the shift, the waveform structure appeared at different positions within the signal. Speech activity was visible near the beginning, middle, and end of the waveform, while some lower-energy regions appeared between these speech segments. The final portion of the shifted signal appeared denser, indicating that high-energy speech regions were moved toward the end of the audio segment.\n\nNo obvious distortion or waveform deformation was observed. The augmentation successfully changed the temporal alignment of the speech while preserving the underlying acoustic content. This demonstrates how time shifting can generate additional training variations without modifying the spoken information itself.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 6: Data Augmentation by Adding Noise\n\n## Objective\n\nThe objective of this task is to increase the robustness of speech data by adding artificial noise to the audio signal. Noise augmentation helps ASR systems perform better in real-world environments where recordings may contain background disturbances.\n\n## Conceptual Understanding\n\nReal-world speech recordings often contain environmental noise such as traffic, fans, room echo, or microphone interference. By adding controlled random noise to clean speech signals, we can simulate challenging recording conditions and create more diverse training data.","metadata":{}},{"cell_type":"code","source":"noise_factor = 0.005\n\nnoise = np.random.randn(len(processing_audio))\n\ny_noisy = processing_audio + noise_factor * noise\n\nplt.figure(figsize=(12,4))\n\nplt.plot(processing_audio[:5000], label=\"Original\", alpha=0.7)\nplt.plot(y_noisy[:5000], label=\"Noisy\", alpha=0.7)\n\nplt.title(\"Original vs Noisy Signal\")\nplt.xlabel(\"Sample Index\")\nplt.ylabel(\"Amplitude\")\n\nplt.legend()\n\nplt.show()\n\nprint(\"Original Min:\", np.min(processing_audio))\nprint(\"Original Max:\", np.max(processing_audio))\n\nprint(\"\\nNoisy Min:\", np.min(y_noisy))\nprint(\"Noisy Max:\", np.max(y_noisy))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:18.734247Z","iopub.execute_input":"2026-06-24T14:54:18.734738Z","iopub.status.idle":"2026-06-24T14:54:19.015843Z","shell.execute_reply.started":"2026-06-24T14:54:18.734704Z","shell.execute_reply":"2026-06-24T14:54:19.014626Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nAfter adding random noise to the speech signal, the waveform became noticeably denser and more irregular compared to the original recording. The overall speech structure remained visible, but small fluctuations appeared throughout the signal due to the injected noise.\n\nThe amplitude range increased slightly from approximately -0.406 to 0.425 in the original signal to approximately -0.416 to 0.436 in the noisy signal. This increase is expected because random noise introduces additional positive and negative amplitude variations.\n\nThe final portion of the waveform appeared slightly taller and denser than before, indicating that noise affected the signal energy across the entire recording. Despite the added disturbances, the underlying speech pattern remained recognizable. This demonstrates how noise augmentation can simulate realistic recording conditions while preserving the original spoken content.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 7: ASR Results Evaluation Using Word Error Rate (WER)\n\n## Objective\n\nThe objective of this task is to evaluate transcription accuracy using the Word Error Rate (WER) metric. WER is the most common evaluation metric used in Automatic Speech Recognition systems and measures the difference between a reference transcript and a predicted transcript.\n\n## Conceptual Understanding\n\nWord Error Rate is calculated using substitutions, insertions, and deletions required to transform a predicted transcript into the correct reference transcript. A lower WER indicates better recognition performance, while a higher WER indicates more transcription errors.","metadata":{}},{"cell_type":"markdown","source":"## Helper function to calculate wer","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef compute_wer(reference, hypothesis):\n\n    ref_words = reference.split()\n    hyp_words = hypothesis.split()\n\n    d = np.zeros(\n        (len(ref_words) + 1, len(hyp_words) + 1),\n        dtype=int\n    )\n\n    for i in range(len(ref_words) + 1):\n        d[i][0] = i\n\n    for j in range(len(hyp_words) + 1):\n        d[0][j] = j\n\n    for i in range(1, len(ref_words) + 1):\n        for j in range(1, len(hyp_words) + 1):\n\n            cost = (\n                0\n                if ref_words[i - 1] == hyp_words[j - 1]\n                else 1\n            )\n\n            d[i][j] = min(\n                d[i - 1][j] + 1,      # deletion\n                d[i][j - 1] + 1,      # insertion\n                d[i - 1][j - 1] + cost  # substitution\n            )\n\n    return d[len(ref_words)][len(hyp_words)] / len(ref_words)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:19.017704Z","iopub.execute_input":"2026-06-24T14:54:19.018165Z","iopub.status.idle":"2026-06-24T14:54:19.027485Z","shell.execute_reply.started":"2026-06-24T14:54:19.018105Z","shell.execute_reply":"2026-06-24T14:54:19.026037Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"reference = \"আজ আমরা বাংলা ভাষার স্বয়ংক্রিয় বক্তৃতা শনাক্তকরণ নিয়ে আলোচনা করব\"\n\nhypothesis = \"আজ আমরা বাংলা ভাষার বক্তৃতা শনাক্তকরণ নিয়ে আলোচনা করব\"\n\nwer = compute_wer(reference, hypothesis)\n\nprint(\"Reference:\")\nprint(reference)\n\nprint(\"\\nHypothesis:\")\nprint(hypothesis)\n\nprint(\"\\nWER =\", round(wer, 4))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:19.029137Z","iopub.execute_input":"2026-06-24T14:54:19.030192Z","iopub.status.idle":"2026-06-24T14:54:19.059815Z","shell.execute_reply.started":"2026-06-24T14:54:19.030131Z","shell.execute_reply":"2026-06-24T14:54:19.058579Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nThe Word Error Rate (WER) for the example transcription was calculated as 0.10, corresponding to a 10% error rate. The predicted transcript differed from the reference transcript by the omission of a single word.\n\nThis experiment demonstrates how even a small transcription mistake contributes directly to the final ASR evaluation score. Since WER is based on insertions, deletions, and substitutions, missing words can significantly impact performance, especially in long-form speech recognition tasks where errors may accumulate over thousands of words.\n\nThe example highlights the importance of accurate transcription and illustrates why WER is the standard evaluation metric used in speech recognition competitions and research studies.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 8: Audio Amplitude Normalization\n\n## Objective\n\nThe objective of this task is to normalize the amplitude of the speech signal so that all values fall within a standardized range. Normalization helps reduce volume differences between recordings and improves consistency during feature extraction.\n\n## Conceptual Understanding\n\nDifferent recordings may have different loudness levels due to microphone settings, speaker distance, or recording environments. Amplitude normalization scales the waveform while preserving its shape, ensuring that signal intensity remains consistent across samples.","metadata":{}},{"cell_type":"code","source":"y_norm = processing_audio / np.max(np.abs(processing_audio))\n\nprint(\"Before Normalization\")\nprint(\"Min:\", np.min(processing_audio))\nprint(\"Max:\", np.max(processing_audio))\n\nprint(\"\\nAfter Normalization\")\nprint(\"Min:\", np.min(y_norm))\nprint(\"Max:\", np.max(y_norm))\n\nplt.figure(figsize=(12,4))\n\nplt.plot(processing_audio[:5000], label=\"Original\")\nplt.plot(y_norm[:5000], label=\"Normalized\")\n\nplt.title(\"Original vs Normalized Signal\")\nplt.xlabel(\"Sample Index\")\nplt.ylabel(\"Amplitude\")\n\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:19.061083Z","iopub.execute_input":"2026-06-24T14:54:19.061401Z","iopub.status.idle":"2026-06-24T14:54:19.298575Z","shell.execute_reply.started":"2026-06-24T14:54:19.061375Z","shell.execute_reply":"2026-06-24T14:54:19.297276Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nAmplitude normalization successfully scaled the speech signal to a standardized range. Before normalization, the waveform amplitudes ranged from approximately -0.406 to 0.425. After normalization, the amplitudes ranged from approximately -0.954 to 1.0.\n\nThe overall waveform structure remained preserved, indicating that normalization changed only the signal scale and not the underlying speech content. Speech activity patterns remained visible throughout the recording, while the amplitude values became more consistent and comparable across the entire signal.\n\nVisual inspection showed that some regions appeared taller and slightly denser after normalization due to the increased scaling of the waveform. This is expected because normalization amplifies the signal relative to its maximum absolute amplitude. Such preprocessing is important in ASR pipelines because it reduces loudness variations between recordings and improves feature extraction consistency.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 9: Silence Stripping Using Voice Activity Detection\n\n## Objective\n\nThe objective of this task is to identify and remove low-energy silent regions from the speech signal. Eliminating silence reduces unnecessary processing and allows feature extraction methods to focus on speech content.\n\n## Conceptual Understanding\n\nSpeech recordings often contain pauses, breathing sounds, and inactive regions before or after speech. Voice Activity Detection (VAD) techniques identify these low-energy segments and remove them while preserving spoken information. This can improve computational efficiency and potentially improve ASR performance.","metadata":{}},{"cell_type":"code","source":"trimmed_audio, index = librosa.effects.trim(\n    processing_audio,\n    top_db=25\n)\n\noriginal_duration = len(processing_audio) / sr\ntrimmed_duration = len(trimmed_audio) / sr\n\nremoved_duration = original_duration - trimmed_duration\n\nprint(\"Original Length (sec):\", round(original_duration, 3))\nprint(\"Trimmed Length (sec):\", round(trimmed_duration, 3))\nprint(\"Silence Removed (sec):\", round(removed_duration, 3))\n\nplt.figure(figsize=(12,4))\n\nlibrosa.display.waveshow(\n    trimmed_audio,\n    sr=sr\n)\n\nplt.title(\"Waveform After Silence Removal\")\nplt.xlabel(\"Time (seconds)\")\nplt.ylabel(\"Amplitude\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:19.300895Z","iopub.execute_input":"2026-06-24T14:54:19.301391Z","iopub.status.idle":"2026-06-24T14:54:19.879592Z","shell.execute_reply.started":"2026-06-24T14:54:19.301349Z","shell.execute_reply":"2026-06-24T14:54:19.878443Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nVoice Activity Detection successfully identified and removed low-energy regions from the speech signal. The original audio segment had a duration of 10.0 seconds, while the trimmed signal had a duration of approximately 9.616 seconds. This indicates that approximately 0.384 seconds of silence or low-energy audio was removed.\n\nVisual inspection of the waveform showed varying speech activity throughout the recording. Lower-amplitude regions were visible during the beginning of the segment, followed by stronger speech activity with higher amplitude peaks. Additional periods of reduced energy appeared later in the recording, including a noticeable low-activity region near the end of the segment before speech activity resumed.\n\nSince less than half a second of audio was removed, the selected segment appears to contain mostly continuous speech with relatively few silent intervals. This suggests that the recording is well-suited for speech recognition tasks and requires only minimal silence removal preprocessing.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 10: Frequency-Domain Audio Visualization\n\n## Objective\n\nThe objective of this task is to analyze the speech signal in the frequency domain using the Fast Fourier Transform (FFT). Frequency-domain analysis helps identify which frequency components contribute most strongly to the speech signal.\n\n## Conceptual Understanding\n\nWhile the time-domain waveform shows how amplitude changes over time, the frequency-domain representation reveals how signal energy is distributed across frequencies. Human speech typically contains most of its energy in lower and middle frequency ranges, making frequency analysis an important component of ASR preprocessing.","metadata":{}},{"cell_type":"code","source":"fft = np.fft.fft(processing_audio)\n\nmagnitude = np.abs(fft[:len(fft)//2])\n\nfreqs = np.fft.fftfreq(\n    len(processing_audio),\n    d=1/sr\n)\n\nfreqs = freqs[:len(freqs)//2]\n\nplt.figure(figsize=(12,4))\n\nplt.plot(freqs, magnitude)\n\nplt.title(\"Frequency Spectrum of Speech Signal\")\nplt.xlabel(\"Frequency (Hz)\")\nplt.ylabel(\"Magnitude\")\n\nplt.xlim(0, 8000)\n\nplt.show()\n\npeak_frequency = freqs[np.argmax(magnitude)]\n\nprint(\"Dominant Frequency (Hz):\", peak_frequency)\nprint(\"Maximum Magnitude:\", np.max(magnitude))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:19.880858Z","iopub.execute_input":"2026-06-24T14:54:19.881216Z","iopub.status.idle":"2026-06-24T14:54:20.138521Z","shell.execute_reply.started":"2026-06-24T14:54:19.881180Z","shell.execute_reply":"2026-06-24T14:54:20.137430Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nThe Fast Fourier Transform (FFT) was used to analyze the speech signal in the frequency domain. The dominant frequency component was observed at approximately 30.8 Hz with a maximum magnitude of approximately 516.48.\n\nThe frequency spectrum showed strong energy concentration in the lower-frequency region, particularly below 1000 Hz. Beyond this range, the magnitude gradually decreased as frequency increased. This behavior is characteristic of human speech signals because most speech energy is concentrated in lower and middle frequencies, while higher-frequency components generally contain less energy.\n\nThe spectrum exhibited a smooth decline toward higher frequencies rather than abrupt fluctuations, indicating that the speech recording contains meaningful acoustic structure rather than random noise. These observations confirm that the dataset possesses the frequency characteristics expected in long-form speech recordings and is appropriate for ASR feature extraction.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 11: Text Label Preprocessing\n\n## Objective\n\nThe objective of this task is to clean and standardize transcript text before it is used in an ASR pipeline. Proper text preprocessing ensures consistency between training labels and model predictions.\n\n## Conceptual Understanding\n\nText data may contain inconsistent spacing, formatting variations, or multiple Unicode representations of the same characters. Unicode NFC normalization converts text into a standardized format, reducing mismatches during evaluation. This is particularly important for Bengali ASR systems because visually identical characters may have different underlying Unicode encodings.","metadata":{}},{"cell_type":"code","source":"import unicodedata\nimport re\n\nsample_file = annotation_files[0]\n\nwith open(\n    os.path.join(annotation_dir, sample_file),\n    \"r\",\n    encoding=\"utf-8\"\n) as f:\n    raw_text = f.read()\n\ndef preprocess_text(text):\n\n    text = unicodedata.normalize(\"NFC\", text)\n\n    text = re.sub(r\"\\s+\", \" \", text)\n\n    return text.strip()\n\nclean_text = preprocess_text(raw_text)\n\nprint(\"Original Character Count:\", len(raw_text))\nprint(\"Processed Character Count:\", len(clean_text))\n\nprint(\"\\nFirst 300 Characters:\\n\")\nprint(clean_text[:300])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.139853Z","iopub.execute_input":"2026-06-24T14:54:20.140358Z","iopub.status.idle":"2026-06-24T14:54:20.157937Z","shell.execute_reply.started":"2026-06-24T14:54:20.140326Z","shell.execute_reply":"2026-06-24T14:54:20.156772Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nThe transcript preprocessing pipeline was applied using Unicode NFC normalization and whitespace standardization. The selected transcript contained 29,326 characters before preprocessing and 29,326 characters after preprocessing.\n\nSince the character count remained unchanged, the transcript appears to be well-structured and already stored in a consistent format. No significant formatting issues, redundant spaces, or encoding irregularities were detected in the selected sample.\n\nThe transcript consists of continuous Bengali text containing narrative speech content. Unicode NFC normalization was still applied because the competition evaluation process performs text comparisons using NFC-normalized strings. Applying this preprocessing step helps prevent potential mismatches caused by different Unicode representations of visually identical Bengali characters.\n\nThe results suggest that the dataset labels are relatively clean and require minimal textual preprocessing before use in an ASR pipeline.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 12: Feature Extraction Using Linear Predictive Coding (LPC)\n\n## Objective\n\nThe objective of this task is to compute Linear Predictive Coding (LPC) coefficients from a speech signal. LPC is a traditional speech processing technique used to model the vocal tract and capture important spectral characteristics of speech.\n\n## Conceptual Understanding\n\nLPC estimates how a speech sample can be predicted from previous samples. The resulting coefficients describe the spectral envelope of the signal and provide a compact representation of speech characteristics. LPC features have historically been used in speech recognition, speaker identification, and speech synthesis systems.","metadata":{}},{"cell_type":"code","source":"frame = processing_audio[:400]\n\nlpc_order = 16\n\nlpc_coeffs = librosa.lpc(\n    frame,\n    order=lpc_order\n)\n\nprint(\"LPC Order:\", lpc_order)\nprint(\"\\nLPC Coefficients:\\n\")\nprint(lpc_coeffs)\n\nplt.figure(figsize=(10,4))\n\nplt.stem(lpc_coeffs)\n\nplt.title(f\"LPC Coefficients (Order {lpc_order})\")\nplt.xlabel(\"Coefficient Index\")\nplt.ylabel(\"Coefficient Value\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.159404Z","iopub.execute_input":"2026-06-24T14:54:20.159781Z","iopub.status.idle":"2026-06-24T14:54:20.360290Z","shell.execute_reply.started":"2026-06-24T14:54:20.159750Z","shell.execute_reply":"2026-06-24T14:54:20.359240Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nLinear Predictive Coding (LPC) coefficients were extracted from a speech frame using an LPC order of 16. The resulting coefficient vector contained 17 values, including the leading coefficient of 1.0.\n\nThe LPC coefficients exhibited alternating positive and negative values, indicating that the model captured complex spectral relationships within the speech signal. Several coefficients had relatively large magnitudes, while others were smaller, suggesting that different prediction terms contribute unequally to modeling the speech waveform.\n\nThe stem plot showed a mixture of long and short coefficient magnitudes rather than a uniform pattern. This behavior is expected because LPC attempts to model the vocal tract characteristics and spectral envelope of speech using a limited set of predictive parameters.\n\nThe extracted coefficients provide a compact mathematical representation of the speech signal and demonstrate how LPC can capture important acoustic information for speech-processing applications.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 13: Data Augmentation Using Pitch Shifting\n\n## Objective\n\nThe objective of this task is to modify the pitch of the speech signal while preserving its duration. Pitch shifting is commonly used as a data augmentation technique to increase speaker variability and improve model robustness.\n\n## Conceptual Understanding\n\nDifferent speakers naturally have different vocal pitch characteristics. Pitch shifting simulates these variations by raising or lowering the perceived frequency of speech without changing the spoken content or recording length. This helps ASR systems generalize better to diverse speakers.","metadata":{}},{"cell_type":"code","source":"n_steps = 2\n\ny_pitch = librosa.effects.pitch_shift(\n    processing_audio,\n    sr=sr,\n    n_steps=n_steps\n)\n\nprint(\"Pitch Shift Applied:\", n_steps, \"semitones\")\n\nprint(\"Original Duration:\",\n      round(len(processing_audio)/sr, 3),\n      \"seconds\")\n\nprint(\"Pitch Shifted Duration:\",\n      round(len(y_pitch)/sr, 3),\n      \"seconds\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.361511Z","iopub.execute_input":"2026-06-24T14:54:20.361814Z","iopub.status.idle":"2026-06-24T14:54:20.461361Z","shell.execute_reply.started":"2026-06-24T14:54:20.361788Z","shell.execute_reply":"2026-06-24T14:54:20.460252Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Audio player to listen to audio","metadata":{}},{"cell_type":"code","source":"from IPython.display import Audio\n\nAudio(y_pitch, rate=sr)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.463014Z","iopub.execute_input":"2026-06-24T14:54:20.463455Z","iopub.status.idle":"2026-06-24T14:54:20.476873Z","shell.execute_reply.started":"2026-06-24T14:54:20.463410Z","shell.execute_reply":"2026-06-24T14:54:20.475785Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nPitch shifting was applied to the speech signal by increasing the pitch by 2 semitones. The original audio duration was 10.0 seconds, and the pitch-shifted audio also maintained a duration of 10.0 seconds.\n\nThis confirms that pitch shifting modified the perceived vocal frequency while preserving the temporal structure of the recording. Unlike speed perturbation, pitch shifting does not compress or expand the signal in time.\n\nBy artificially creating speech with different vocal characteristics, pitch shifting can increase speaker diversity within a dataset and improve the robustness of speech recognition systems. The augmentation successfully generated an alternative version of the speech signal while retaining the original spoken content and duration.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 14: Data Augmentation Using Speed Perturbation\n\n## Objective\n\nThe objective of this task is to modify the playback speed of the speech signal. Speed perturbation is a widely used augmentation technique that creates variations in speaking rate while preserving the linguistic content.\n\n## Conceptual Understanding\n\nDifferent speakers naturally speak at different speeds. Speed perturbation simulates these variations by compressing or stretching the audio signal in time. This helps ASR systems become more robust to changes in speaking rate and pronunciation timing.","metadata":{}},{"cell_type":"code","source":"speed_factor = 1.2\n\ny_fast = librosa.effects.time_stretch(\n    processing_audio,\n    rate=speed_factor\n)\n\nprint(\"Speed Factor:\", speed_factor)\n\nprint(\"Original Duration:\",\n      round(len(processing_audio)/sr, 3),\n      \"seconds\")\n\nprint(\"Modified Duration:\",\n      round(len(y_fast)/sr, 3),\n      \"seconds\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.478376Z","iopub.execute_input":"2026-06-24T14:54:20.478885Z","iopub.status.idle":"2026-06-24T14:54:20.559024Z","shell.execute_reply.started":"2026-06-24T14:54:20.478843Z","shell.execute_reply":"2026-06-24T14:54:20.557798Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Graph","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,4))\n\nplt.plot(processing_audio[:5000], label=\"Original\")\n\nplt.plot(y_fast[:5000], label=\"Speed Perturbed\")\n\nplt.title(\"Original vs Speed Perturbed Signal\")\n\nplt.xlabel(\"Sample Index\")\nplt.ylabel(\"Amplitude\")\n\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.560376Z","iopub.execute_input":"2026-06-24T14:54:20.560713Z","iopub.status.idle":"2026-06-24T14:54:20.826888Z","shell.execute_reply.started":"2026-06-24T14:54:20.560679Z","shell.execute_reply":"2026-06-24T14:54:20.825263Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nSpeed perturbation was applied using a speed factor of 1.2. The original audio duration was 10.0 seconds, while the modified audio duration decreased to approximately 8.333 seconds.\n\nVisual inspection of the waveform showed that the speech signal became more compressed and denser during the earlier portion of the recording. As the waveform progressed, the signal structure appeared more compact compared to the original version due to the increased playback speed. The overlapping waveform visualization also showed greater separation between the original and modified signals as time progressed.\n\nThe reduction in duration confirms that the speech was played faster while preserving the underlying linguistic content. Speed perturbation is a valuable augmentation technique because it exposes ASR systems to different speaking rates, helping improve recognition performance across speakers with varying speech tempos.\n","metadata":{}},{"cell_type":"markdown","source":"# Task 15: Audio Downsampling Transformation\n\n## Objective\n\nThe objective of this task is to reduce the sampling rate of the speech signal. Downsampling decreases the number of samples used to represent the audio and can reduce storage and computational requirements.\n\n## Conceptual Understanding\n\nThe sampling rate determines how many audio samples are recorded per second. A higher sampling rate captures more frequency information, while a lower sampling rate reduces data size but may remove high-frequency content. Many ASR systems use standardized sampling rates to ensure consistency across datasets.","metadata":{}},{"cell_type":"code","source":"target_sr = 8000\n\ndownsampled_audio = librosa.resample(\n    processing_audio,\n    orig_sr=sr,\n    target_sr=target_sr\n)\n\nprint(\"Original Sample Rate:\", sr, \"Hz\")\nprint(\"Target Sample Rate:\", target_sr, \"Hz\")\n\nprint(\"\\nOriginal Samples:\", len(processing_audio))\nprint(\"Downsampled Samples:\", len(downsampled_audio))\n\nprint(\"\\nOriginal Duration:\",\n      round(len(processing_audio)/sr, 3),\n      \"seconds\")\n\nprint(\"Downsampled Duration:\",\n      round(len(downsampled_audio)/target_sr, 3),\n      \"seconds\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.828415Z","iopub.execute_input":"2026-06-24T14:54:20.828792Z","iopub.status.idle":"2026-06-24T14:54:20.839327Z","shell.execute_reply.started":"2026-06-24T14:54:20.828763Z","shell.execute_reply":"2026-06-24T14:54:20.838107Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Graph","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,4))\n\nplt.plot(processing_audio[:5000], label=\"Original\")\n\nplt.plot(downsampled_audio[:2500], label=\"Downsampled\")\n\nplt.title(\"Original vs Downsampled Audio\")\n\nplt.xlabel(\"Sample Index\")\nplt.ylabel(\"Amplitude\")\n\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-24T14:54:20.840680Z","iopub.execute_input":"2026-06-24T14:54:20.841030Z","iopub.status.idle":"2026-06-24T14:54:21.120482Z","shell.execute_reply.started":"2026-06-24T14:54:20.840994Z","shell.execute_reply":"2026-06-24T14:54:21.119380Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Findings\n\nThe speech signal was successfully downsampled from 16,000 Hz to 8,000 Hz. As expected, the total number of samples decreased from 160,000 to 80,000, representing a 50% reduction in data size.\n\nDespite the reduction in sample count, the duration of the recording remained unchanged at 10.0 seconds. This confirms that downsampling changes the temporal resolution of the signal rather than its playback length.\n\nVisual inspection showed that the downsampled waveform contained fewer samples and appeared less dense than the original signal. The original waveform occupied a larger portion of the graph, while the downsampled version represented the same audio information using fewer points. This demonstrates how downsampling can reduce storage and computational requirements while preserving the overall speech content.\n\nThe dataset originally uses a sampling rate of 16 kHz, which is already a standard sampling rate for many Automatic Speech Recognition systems. Therefore, additional downsampling may not be necessary in a practical ASR pipeline unless computational efficiency is a priority.\n","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\n\nThis case study examined a Bengali long-form speech recognition dataset through a complete preprocessing and feature extraction pipeline.\n\nThe dataset consists of long-duration recordings with consistent 16 kHz sampling rates and substantial transcript lengths. Analysis of both audio and text data showed that the dataset is relatively clean and well-structured, requiring only limited preprocessing.\n\nTime-domain and frequency-domain visualizations demonstrated that speech energy is concentrated primarily in lower frequencies, which is consistent with expected human speech characteristics. Feature extraction techniques such as Log-Mel Spectrograms and LPC successfully captured important acoustic information.\n\nData augmentation techniques including time shifting, noise addition, pitch shifting, and speed perturbation generated realistic variations of the speech signal while preserving linguistic content. Silence stripping removed only a small amount of inactive audio, suggesting that the recordings contain mostly continuous speech.\n\nOverall, the dataset appears suitable for Automatic Speech Recognition research and development. The most important preprocessing steps identified during this study are amplitude normalization, Unicode text normalization, feature extraction, and controlled data augmentation for improved model robustness.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}