{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Dataset Exploration\n\n## Engineering Objective\n\nThe objective of this exploratory analysis is to understand the overall structure, scale, and characteristics of the Quran ASR dataset before performing speech preprocessing and feature extraction tasks. This analysis helps identify the available audio recordings, transcription labels, speaker diversity, and dataset organization required for building an Automatic Speech Recognition (ASR) pipeline.\n\n## Conceptual Understanding\n\nDataset exploration is a fundamental step in speech processing systems. It provides insights into the number of speech samples, speaker distribution, transcript availability, audio formats, and corpus complexity. Understanding these properties enables informed decisions regarding preprocessing techniques, feature engineering methods, and augmentation strategies.\n\nFor ASR applications, exploratory analysis assists in identifying:\n- Dataset size and scalability\n- Speaker variability\n- Recording formats\n- Text transcription characteristics\n- Potential biases in speech duration and speaker representation\n\n## Dataset Findings\n\nThe Quran ASR dataset contains:\n\n- **71,054 audio recordings** in the training set.\n- **15,000 audio recordings** in the testing set.\n- Audio recordings are stored in **MP3 format**.\n- A transcription file named **train_transcriptions.csv** is provided.\n- The transcription dataset consists of three attributes:\n  - `audio_id`\n  - `speaker_id`\n  - `text`\n- Multiple speakers are represented within the corpus.\n- The dataset is designed for large-scale Arabic Quranic speech recognition experiments.\n\n## Preliminary Observations\n\nInitial exploration indicates that the dataset is sufficiently large for speech recognition research and contains diverse speaker samples. Further analysis will be conducted to investigate speech durations, transcript lengths, amplitude distributions, and spectral properties of the recordings.","metadata":{}},{"cell_type":"code","source":"import os\n\nfor root, dirs, files in os.walk('/kaggle/input'):\n    print(\"folder:\",root)\n    print(\"sub folders:\",len(dirs))\n    print(\"files:\",len(files))\n    print(\"-\"*60)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-28T18:48:41.417636Z","iopub.execute_input":"2026-06-28T18:48:41.417910Z","iopub.status.idle":"2026-06-28T18:49:14.841794Z","shell.execute_reply.started":"2026-06-28T18:48:41.417886Z","shell.execute_reply":"2026-06-28T18:49:14.840794Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\nbase=\"/kaggle/input/competitions/quran-asr-challenge\"\n\nprint(os.listdir(base))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-28T18:55:07.135313Z","iopub.execute_input":"2026-06-28T18:55:07.136122Z","iopub.status.idle":"2026-06-28T18:55:07.142105Z","shell.execute_reply.started":"2026-06-28T18:55:07.136088Z","shell.execute_reply":"2026-06-28T18:55:07.141072Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\n\ndf = pd.read_csv(\n'/kaggle/input/competitions/quran-asr-challenge/train_transcriptions.csv'\n)\n\ndf.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-28T19:01:19.184864Z","iopub.execute_input":"2026-06-28T19:01:19.185364Z","iopub.status.idle":"2026-06-28T19:01:19.420057Z","shell.execute_reply.started":"2026-06-28T19:01:19.185318Z","shell.execute_reply":"2026-06-28T19:01:19.419187Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-28T19:01:39.954106Z","iopub.execute_input":"2026-06-28T19:01:39.954621Z","iopub.status.idle":"2026-06-28T19:01:39.961665Z","shell.execute_reply.started":"2026-06-28T19:01:39.954589Z","shell.execute_reply":"2026-06-28T19:01:39.960474Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\n\ndf= pd.read_csv ('/kaggle/input/competitions/quran-asr-challenge/train_transcriptions.csv')\n\nprint(df.head())\n\nprint()\n\nprint(\"Shape:\",df.shape)\n\nprint()\n\nprint(\"Unique speakers:\",\n     df['speaker_id'].nunique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-28T19:08:43.787013Z","iopub.execute_input":"2026-06-28T19:08:43.787959Z","iopub.status.idle":"2026-06-28T19:08:44.040584Z","shell.execute_reply.started":"2026-06-28T19:08:43.787905Z","shell.execute_reply":"2026-06-28T19:08:44.039334Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"files = os.listdir('/kaggle/input/competitions/quran-asr-challenge/train_set')\n\nprint(files[:10])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-28T19:09:34.955233Z","iopub.execute_input":"2026-06-28T19:09:34.955563Z","iopub.status.idle":"2026-06-28T19:09:35.418725Z","shell.execute_reply.started":"2026-06-28T19:09:34.955536Z","shell.execute_reply":"2026-06-28T19:09:35.417571Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 1: Input Libraries & Audio Loading\n\n## Engineering Objective\n\nThe objective of this task is to load a speech recording from the Quran ASR dataset and analyze its fundamental characteristics including sampling frequency, signal duration, and amplitude distribution. Understanding these properties is essential before applying preprocessing and feature extraction techniques in an Automatic Speech Recognition (ASR) pipeline.\n\n## Conceptual Understanding\n\nSpeech recordings are represented digitally as sequences of amplitude samples collected at a fixed sampling frequency. Loading the audio signal enables inspection of recording quality, duration, and signal intensity which are important for subsequent speech analysis tasks.\n\n## Dataset Findings\n\nA Quranic recitation sample was successfully loaded from the training dataset.\n\nThe selected recording is stored in MP3 format and contains a total of 111,744 samples.\n\nThe audio duration is approximately 5.07 seconds.\n\nThe speech signal is sampled at 22,050 Hz.\n\nThe minimum observed amplitude is -0.7935 while the maximum amplitude is 0.9250.\n\nThese observations suggest that the speech recording has been normalized to a bounded amplitude range and contains sufficient temporal information for ASR preprocessing.\n\n### Observations\n\n• Audio format : MP3\n\n• Sampling Rate : 22,050 Hz\n\n• Total Samples : 111,744\n\n• Duration : 5.07 seconds\n\n• Minimum Amplitude : -0.7935\n\n• Maximum Amplitude : 0.9250","metadata":{}},{"cell_type":"code","source":"import librosa\nimport os\n\ntrain_path= '/kaggle/input/competitions/quran-asr-challenge/train_set'\n\nsample_file= os.listdir(train_path)[0]\n\naudio_path= os.path.join(train_path, sample_file)\n\nsignal, sr = librosa.load(audio_path, sr=None)\n\nprint(\"File Name :\", sample_file)\n\nprint(\"sampling Rate:\", sr)\n\nprint(\"Total Samples :\", len(signal))\n\nprint(\"Duration:\", len(signal)/sr)\n\nprint(\"minimum amplitude:\", signal.min())\n\nprint(\"maximum amplitude:\", signal.max())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T03:36:15.704927Z","iopub.execute_input":"2026-06-29T03:36:15.705993Z","iopub.status.idle":"2026-06-29T03:36:40.008355Z","shell.execute_reply.started":"2026-06-29T03:36:15.705956Z","shell.execute_reply":"2026-06-29T03:36:40.007439Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 2: Time-Domain Audio Visualization\n\n## Engineering Objective\n\nThe objective of this task is to visualize the temporal structure of speech signals in order to identify speech activity, silence intervals, and amplitude fluctuations over time.\n\n## Conceptual Understanding\n\nTime-domain visualizations represent speech amplitude as a function of time. Three-dimensional representations provide additional insight into signal dynamics and energy distributions across the recording.\n\nSpeech signals typically exhibit alternating regions of high energy corresponding to spoken content and low energy regions corresponding to pauses or silence.\n\n## Dataset Findings\n\nThe selected Quranic recitation demonstrates several high-energy regions associated with active speech segments.\n\nLower-amplitude intervals indicate brief pauses between recitations.\n\nThe speech waveform remains bounded within the amplitude range observed in Task 1.\n\nThe temporal distribution suggests a continuous recitation style with limited silence periods.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom mpl_toolkits.mplot3d import Axes3D\nimport numpy as np\n\ntime = np.arange(len(signal))/sr\n\nfig = plt.figure(figsize=(14,6))\n\nax = fig.add_subplot(111, projection='3d')\n\nax.plot(\n        time,\n        signal,\n        np.zeros_like(signal),\n        linewidth=1\n)\n\nax.set_xlabel('Time (s)')\n\nax.set_ylabel('Amplitude')\n\nax.set_zlabel('Depth')\n\nax.set_title('3D Time Domain Speech Visualization')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-28T20:00:02.965011Z","iopub.execute_input":"2026-06-28T20:00:02.966487Z","iopub.status.idle":"2026-06-28T20:00:03.395374Z","shell.execute_reply.started":"2026-06-28T20:00:02.966444Z","shell.execute_reply":"2026-06-28T20:00:03.394301Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Speech activity is concentrated between 0.5 and 4.8 seconds.\n\nSeveral low-energy intervals are visible.\n\nAmplitude peaks approach ±0.9.\n\nNo clipping artifacts are observed.\n\nThe recording appears clean with limited background noise.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 3 : Windowing & Framing Implementation\n\n## Engineering Objective\n\nThe objective of this task is to segment the continuous speech signal into smaller overlapping frames suitable for spectral analysis and feature extraction. Windowing and framing are essential preprocessing steps in ASR systems because speech signals are considered quasi-stationary over short intervals.\n\n## Conceptual Understanding\n\nSpeech signals are inherently non-stationary over long durations but can be approximated as stationary over small time windows.\n\nA typical ASR system uses:\n\n• Frame Size = 25 ms\n\n• Hop Length = 10 ms\n\nWindowing helps reduce spectral leakage and preserves local signal characteristics.\n\n## Dataset Findings\n\nThe selected Quranic recitation was divided into short overlapping frames using a Hamming window.\n\nThe framing process generated multiple segments suitable for feature extraction.\n\nFrame overlap preserves temporal continuity and enables better speech representation.\n\n### Observations\n\n• Frame Length : 551 samples\n\n• Hop Length : 220 samples\n\n• Total Frames Generated : 506\n\n• Window Type : Hamming\n\n• Frame Duration : 25 ms\n\n• Frame Shift : 10 ms","metadata":{}},{"cell_type":"code","source":"import librosa\nimport numpy as np \n\nframe_length =int(0.025*sr)\nhop_length= int(0.010 * sr)\n\nframes = librosa.util.frame(\n    signal,\n    frame_length=frame_length,\n    hop_length= hop_length\n)\n\nprint (\"frame length:\", frame_length)\nprint (\"Hop length:\", hop_length )\nprint (\"Total Frames:\", frames.shape[1])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T03:36:47.053249Z","iopub.execute_input":"2026-06-29T03:36:47.053735Z","iopub.status.idle":"2026-06-29T03:36:47.060796Z","shell.execute_reply.started":"2026-06-29T03:36:47.053703Z","shell.execute_reply":"2026-06-29T03:36:47.059824Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nplt.figure(figsize=(12,5))\n\nfor i in range(5):\n    plt.plot(frames[:,i])\n\nplt.title(\"first five speech frames\")\n\nplt.xlabel(\"samples\")\n\nplt.ylabel(\"amplitude\")\n\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T03:37:53.641074Z","iopub.execute_input":"2026-06-29T03:37:53.641402Z","iopub.status.idle":"2026-06-29T03:37:53.900724Z","shell.execute_reply.started":"2026-06-29T03:37:53.641376Z","shell.execute_reply":"2026-06-29T03:37:53.899631Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 4 : Feature Extraction (Log-Mel Spectrogram)\n\n## Engineering Objective\n\nThe objective of this task is to extract robust acoustic features using the Log-Mel Spectrogram representation. These features are widely used in ASR systems due to their similarity to human auditory perception.\n\n## Conceptual Understanding\n\nThe Mel scale models the frequency sensitivity of the human ear.\n\nThe logarithmic transformation compresses the dynamic range of speech signals.\n\nLog-Mel Spectrograms provide discriminative information for speech recognition systems.\n\n## Dataset Findings\n\nThe Log-Mel representation highlights frequency regions containing significant speech energy.\n\nMost energy is concentrated in lower frequency bands.\n\nDistinct spectral patterns corresponding to Quranic recitation are visible.\n\n### Observations\n\n• Number of Mel Bands : 128\n\n• Sampling Rate : 22050 Hz\n\n• Energy concentration observed below approximately 4 kHz","metadata":{}},{"cell_type":"code","source":"import librosa.display\n\nmel = librosa.feature.melspectrogram(\n        y=signal,\n        sr=sr,\n        n_mels=128\n)\n\nlog_mel = librosa.power_to_db(\n            mel,\n            ref=np.max\n)\n\nplt.figure(figsize=(12,6))\n\nlibrosa.display.specshow(\n        log_mel,\n        sr=sr,\n        x_axis='time',\n        y_axis='mel'\n)\n\nplt.colorbar(format='%+2.0f dB')\n\nplt.title(\"Log-Mel Spectrogram\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T03:37:58.221092Z","iopub.execute_input":"2026-06-29T03:37:58.221436Z","iopub.status.idle":"2026-06-29T03:37:58.536170Z","shell.execute_reply.started":"2026-06-29T03:37:58.221407Z","shell.execute_reply":"2026-06-29T03:37:58.535229Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Speech energy is predominantly concentrated in lower Mel frequency bands.\n\nSeveral high-intensity regions correspond to voiced speech segments.\n\nThe recitation exhibits clear harmonic structures indicative of continuous vocal activity.","metadata":{}},{"cell_type":"markdown","source":"# Task 5 : Data Augmentation (Time Shifting)\n\n## Engineering Objective\n\nThe objective of this task is to increase speech variability through temporal shifting. Data augmentation improves model robustness and generalization capabilities.\n\n## Conceptual Understanding\n\nTime shifting modifies the starting position of speech without altering the overall signal characteristics.\n\nThis augmentation simulates slight variations in speech onset timing encountered in practical ASR scenarios.\n\n## Dataset Findings\n\nThe selected Quranic recitation was shifted in time while preserving its original duration.\n\nNo information loss was observed.\n\nThe augmentation introduces temporal diversity into the dataset.\n\n### Observations\n\n• Shift Amount : 1000 samples\n\n• Duration Preserved : Yes\n\n• Signal Characteristics : Maintained","metadata":{}},{"cell_type":"code","source":"shift = 1000\n\nshifted_signal = np.roll(\n                    signal,\n                    shift\n)\n\nplt.figure(figsize=(12,4))\n\nplt.plot(signal)\n\nplt.title(\"Original Signal\")\n\nplt.show()\n\nplt.figure(figsize=(12,4))\n\nplt.plot(shifted_signal)\n\nplt.title(\"Time Shifted Signal\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T03:38:03.268889Z","iopub.execute_input":"2026-06-29T03:38:03.269424Z","iopub.status.idle":"2026-06-29T03:38:03.840458Z","shell.execute_reply.started":"2026-06-29T03:38:03.269392Z","shell.execute_reply":"2026-06-29T03:38:03.839334Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 6 : Data Augmentation (Adding Noise)\n\n## Engineering Objective\n\nThe objective of this task is to improve ASR robustness by introducing background noise into speech recordings.\n\n## Conceptual Understanding\n\nNoise augmentation simulates real-world recording conditions and helps speech recognition systems generalize better under noisy environments.\n\n## Dataset Findings\n\nGaussian noise was added to the Quranic recitation sample. The augmented signal maintains intelligibility while exhibiting increased signal variability.\n\n### Observations\n\n• Noise Type : Gaussian\n\n• Noise Factor : 0.005\n\n• Signal Duration : Preserved\n\n• Speech Intelligibility : Maintained","metadata":{}},{"cell_type":"code","source":"noise = np.random.randn(len(signal))\n\nnoise_factor = 0.005\n\nnoisy_signal = signal + noise_factor * noise\n\nplt.figure(figsize=(12,4))\n\nplt.plot(noisy_signal)\n\nplt.title(\"Speech Signal with Added Noise\")\n\nplt.xlabel(\"Samples\")\n\nplt.ylabel(\"Amplitude\")\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T03:38:07.707937Z","iopub.execute_input":"2026-06-29T03:38:07.708968Z","iopub.status.idle":"2026-06-29T03:38:08.218603Z","shell.execute_reply.started":"2026-06-29T03:38:07.708922Z","shell.execute_reply":"2026-06-29T03:38:08.217509Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Task 7 : Word Error Rate Evaluation\n\n## Engineering Objective\n\nEvaluate ASR transcription quality using Word Error Rate.\n\n## Conceptual Understanding\n\nWER measures transcription accuracy by computing insertions, deletions, and substitutions.\n\nWER=(S+D+I)/N\n\n## Dataset Findings\n\nWER was computed between reference and predicted transcriptions.\n\nLower WER values indicate better recognition performance.","metadata":{}},{"cell_type":"code","source":"!pip install jiwer","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T03:39:05.518800Z","iopub.execute_input":"2026-06-29T03:39:05.519472Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from jiwer import wer\n\nreference = \"بسم الله الرحمن الرحيم\"\n\nprediction = \"بسم الله الرحمن الرحيم\"\n\nscore = wer(reference,prediction)\n\nprint(\"WER:\",score)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}