{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <center style=\"font-family: consolas; font-size: 28px; font-weight: bold;\"> Bengali.AI Speech Recognition : Exploratory Data Analysis and Basic Preprocessing</center>\n<center> <b> Recognize Bengali speech from out-of-distribution audio recordings </b> </center>","metadata":{}},{"cell_type":"markdown","source":"## Task Description\n- To recognize Bengali speech from out-of-distribution audio recordings\n- Out-of-distribution means that the data our model will be tested on, is little different. Which is challenging because in such cases the model tends to perform poor.\n- Out-of-distribution recordings can include variations in accent, background noise, recording conditions, speaking styles, or dialects that were not sufficiently represented in the training data.\n\n## Dataset Description\n- Size of the data - **`26GB`**\n- Train Data is Massively Crowdsourced (MaCro) Bengali speech dataset. Roughly **1,200 hours** of audio recordings from ~**24,000 people** from India and Bangladesh. \n- The test set contains samples from **17 different domains** that are not present in training, which is actually the validation set. **Actual test set is hidden for the sake of contest**. The full test set contains about **20 hours of speech** in almost **8000 MP3** audio files. All of the files in the test set are **encoded at a sample rate of 32k, a bit rate of 48k, in one channel**.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"padding:10px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; \"> 🗊 Table of Contents</center></div>\n\n1.  [Imports](#1)\n1.  [What is this competition about?](#2)\n1.  [Peeking into the dataset](#3)\n1.  [Checking single instances](#4)\n1.  [Text EDA](#5)\n1. [Audio EDA](#6)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; \"> Imports and initializations </center>\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport soundfile as sf\nfrom pydub import AudioSegment\nimport IPython.display as ipd\nfrom collections import Counter\nimport os\nfrom wordcloud import WordCloud, STOPWORDS, ImageColorGenerator","metadata":{"execution":{"iopub.status.busy":"2023-08-07T07:45:54.300217Z","iopub.execute_input":"2023-08-07T07:45:54.300589Z","iopub.status.idle":"2023-08-07T07:45:55.917844Z","shell.execute_reply.started":"2023-08-07T07:45:54.300559Z","shell.execute_reply":"2023-08-07T07:45:55.916560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show longer texts\npd.options.display.max_colwidth = 100","metadata":{"execution":{"iopub.status.busy":"2023-07-30T07:58:59.520577Z","iopub.execute_input":"2023-07-30T07:58:59.521014Z","iopub.status.idle":"2023-07-30T07:58:59.527992Z","shell.execute_reply.started":"2023-07-30T07:58:59.520980Z","shell.execute_reply":"2023-07-30T07:58:59.526720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Color palette from\n# https://www.heavy.ai/blog/12-color-palettes-for-telling-better-stories-with-your-data\n\nsns.set()\n\ndef hex_to_rgb(hex_value):\n  h = hex_value.lstrip('#')\n  return tuple(int(h[i:i + 2], 16) / 255.0 for i in (0, 2, 4))\n\ncolor_palette = [\"#ea5545\", \"#f46a9b\", \"#ef9b20\", \n                 \"#edbf33\", \"#ede15b\", \"#bdcf32\", \n                 \"#87bc45\", \"#27aeef\", \"#b33dc6\"]\n\nrgb_colors = list(map(hex_to_rgb, color_palette))\n\nsns.palplot(rgb_colors)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T06:19:30.316989Z","iopub.execute_input":"2023-08-01T06:19:30.317503Z","iopub.status.idle":"2023-08-01T06:19:30.537181Z","shell.execute_reply.started":"2023-08-01T06:19:30.317462Z","shell.execute_reply":"2023-08-01T06:19:30.535497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold; \">Peeking into the dataset</center>\n","metadata":{}},{"cell_type":"markdown","source":"## Folder Structure","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input/bengaliai-speech","metadata":{"execution":{"iopub.status.busy":"2023-08-01T06:19:30.539124Z","iopub.execute_input":"2023-08-01T06:19:30.539550Z","iopub.status.idle":"2023-08-01T06:19:31.655128Z","shell.execute_reply.started":"2023-08-01T06:19:30.539510Z","shell.execute_reply":"2023-08-01T06:19:31.653316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**`train_mp3s/`** -  The training set, comprising several thousand recordings in **MP3 format**.\n\n**`test_mp3s/`** -  The test set, comprising spontaneous speech recordings from **18 domains**, **17 of which are out-of-distribution with respect to the training set**. There may be domains in the private test set that are **not in the public test set**.\n\n**`examples/`** An **example recording** for **each test set domain**. You may find these example recordings **helpful for creating models robust to domain variation**. These are representative recordings and **none of them are present in the test set**.\n\n**`train.csv`** -  Sentence labels for the training set.","metadata":{}},{"cell_type":"markdown","source":"## CSV Details","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(\"/kaggle/input/bengaliai-speech/train.csv\")\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-07T07:59:33.972474Z","iopub.execute_input":"2023-08-07T07:59:33.974095Z","iopub.status.idle":"2023-08-07T07:59:39.207648Z","shell.execute_reply.started":"2023-08-07T07:59:33.974041Z","shell.execute_reply":"2023-08-07T07:59:39.206033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" - `id` -  A unique identifier for this instance. **Corresponds to the file {id}.mp3 in train/**.\n - `sentence` -  A **plain-text transcription of the recording**. Your goal is to **predict these sentences for each recording in the test set**.\n - `split` -  Whether **train or valid**. The annotations in the **valid** split have been **manually reviewed and corrected**, while the annotations in the **train** split have only been **algorithmically cleaned**. The **valid** samples will generally have **higher quality annotations** than the train samples, but are otherwise **drawn from the same distribution**.","metadata":{}},{"cell_type":"code","source":"df.split.unique()","metadata":{"execution":{"iopub.status.busy":"2023-07-31T06:46:03.685733Z","iopub.execute_input":"2023-07-31T06:46:03.686383Z","iopub.status.idle":"2023-07-31T06:46:03.763567Z","shell.execute_reply.started":"2023-07-31T06:46:03.686346Z","shell.execute_reply":"2023-07-31T06:46:03.762438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many train and how many validation samples?","metadata":{}},{"cell_type":"code","source":"n_train_samples = sum(df[\"split\"]==\"train\")\nn_valid_samples = sum(df[\"split\"]==\"valid\")\nprint(f\"Total training samples : \",n_train_samples)\nprint(f\"Total validation samples : \",n_valid_samples)\nprint(\"Validation/Train ratio : \",n_valid_samples/n_train_samples)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T06:46:06.526113Z","iopub.execute_input":"2023-07-31T06:46:06.526515Z","iopub.status.idle":"2023-07-31T06:46:06.846742Z","shell.execute_reply.started":"2023-07-31T06:46:06.526486Z","shell.execute_reply":"2023-07-31T06:46:06.845646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Okay so validation set is very small compared to training set. We might need to use additional data from the train set as validation set.","metadata":{}},{"cell_type":"code","source":"plt.bar([\"train\",\"valid\"],[n_train_samples,n_valid_samples],color = [color_palette[6], color_palette[0]])","metadata":{"execution":{"iopub.status.busy":"2023-07-31T06:47:07.248162Z","iopub.execute_input":"2023-07-31T06:47:07.248520Z","iopub.status.idle":"2023-07-31T06:47:07.455794Z","shell.execute_reply.started":"2023-07-31T06:47:07.248493Z","shell.execute_reply":"2023-07-31T06:47:07.454723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Random audios with transcriptions","metadata":{}},{"cell_type":"code","source":"root_path = \"/kaggle/input/bengaliai-speech/train_mp3s\"\n\nrandom_idx = np.random.randint(0, len(df), 10)\n\nfor idx in random_idx:\n    mp3_path = os.path.join(root_path,df['id'].iloc[idx])+ \".mp3\"\n    text = df['sentence'].iloc[idx]\n    display(AudioSegment.from_file(mp3_path))\n    print(f'Transcription {idx}: ',text)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-28T10:41:52.031340Z","iopub.execute_input":"2023-07-28T10:41:52.031751Z","iopub.status.idle":"2023-07-28T10:41:55.108502Z","shell.execute_reply.started":"2023-07-28T10:41:52.031705Z","shell.execute_reply":"2023-07-28T10:41:55.107136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From these audios it's quite evident that the audios has a lot of varieties, in domains, in audio quality and other forms. Feel free to change the idxs and listen to different audios","metadata":{}},{"cell_type":"markdown","source":"For the test set, we'll have data from 17 domains. They have provided a sample for each of them. Let's hear some of them.","metadata":{}},{"cell_type":"code","source":"mp3_path = \"/kaggle/input/bengaliai-speech/examples/Cartoon.wav\"\nprint(\"Cartoon Audio\")\ndisplay(AudioSegment.from_file(mp3_path))\n\nmp3_path = \"/kaggle/input/bengaliai-speech/examples/Poem Recital.wav\"\nprint(\"Poem recital \")\ndisplay(AudioSegment.from_file(mp3_path))\n","metadata":{"execution":{"iopub.status.busy":"2023-07-18T17:03:07.90867Z","iopub.execute_input":"2023-07-18T17:03:07.909225Z","iopub.status.idle":"2023-07-18T17:03:09.784949Z","shell.execute_reply.started":"2023-07-18T17:03:07.909183Z","shell.execute_reply":"2023-07-18T17:03:09.783659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This Cartoon is probably from the popular Cartoon show \"Gopal Bhar\" and the poem was written by the National Poet of Bangladesh, \"Kazi Nazrul Islam\"","metadata":{}},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold;\">Checking single instance</center>\n","metadata":{}},{"cell_type":"markdown","source":"### Using audio id","metadata":{}},{"cell_type":"code","source":"audio_id = '0000d8b5ce2b'\npath = f\"/kaggle/input/bengaliai-speech/train_mp3s/{audio_id}.mp3\"\ndisplay(AudioSegment.from_file(path))\nprint(df[df['id'] == audio_id]['sentence'])","metadata":{"execution":{"iopub.status.busy":"2023-07-28T09:06:17.886149Z","iopub.execute_input":"2023-07-28T09:06:17.886644Z","iopub.status.idle":"2023-07-28T09:06:18.358991Z","shell.execute_reply.started":"2023-07-28T09:06:17.886594Z","shell.execute_reply":"2023-07-28T09:06:18.357430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Using dataframe iloc","metadata":{}},{"cell_type":"code","source":"dataframe_id = 2\naudio_id = df.iloc[dataframe_id]['id']\npath = f\"/kaggle/input/bengaliai-speech/train_mp3s/{audio_id}.mp3\"\ndisplay(AudioSegment.from_file(path))\nprint(df.iloc[dataframe_id]['sentence'])","metadata":{"execution":{"iopub.status.busy":"2023-07-28T09:13:19.020403Z","iopub.execute_input":"2023-07-28T09:13:19.020810Z","iopub.status.idle":"2023-07-28T09:13:19.319006Z","shell.execute_reply.started":"2023-07-28T09:13:19.020772Z","shell.execute_reply":"2023-07-28T09:13:19.317966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold;\"> Text EDA </center>\n","metadata":{}},{"cell_type":"code","source":"print(df.shape)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-30T07:59:59.512714Z","iopub.execute_input":"2023-07-30T07:59:59.513237Z","iopub.status.idle":"2023-07-30T07:59:59.529374Z","shell.execute_reply.started":"2023-07-30T07:59:59.513199Z","shell.execute_reply":"2023-07-30T07:59:59.527764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A total of 963636 sentences. But how many unique sentences?","metadata":{}},{"cell_type":"code","source":"print(\"Total sentences :\",len(df))\nprint(\"Total unique sentences : \",df.sentence.nunique())\nprint(\"Percentage of unique sentences ; \",df.sentence.nunique()/len(df))","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:00:04.598357Z","iopub.execute_input":"2023-07-30T08:00:04.599566Z","iopub.status.idle":"2023-07-30T08:00:06.051982Z","shell.execute_reply.started":"2023-07-30T08:00:04.599514Z","shell.execute_reply":"2023-07-30T08:00:06.050262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's interesting. 50% of the overall sentences are actually unique. So multiple of them has been repeated in the dataset. Let's look at the most frequent 10 sentences","metadata":{}},{"cell_type":"code","source":"print(\"Most frequent sentences in the dataset \\n\")\ndf.sentence.value_counts()[:10]","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:00:23.109545Z","iopub.execute_input":"2023-07-30T08:00:23.110076Z","iopub.status.idle":"2023-07-30T08:00:24.326264Z","shell.execute_reply.started":"2023-07-30T08:00:23.110035Z","shell.execute_reply":"2023-07-30T08:00:24.324644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Let's look at the sentence length distribuition**","metadata":{}},{"cell_type":"code","source":"x = df.sentence.apply(lambda x: len(x))","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:01:15.885517Z","iopub.execute_input":"2023-07-30T08:01:15.886032Z","iopub.status.idle":"2023-07-30T08:01:16.814598Z","shell.execute_reply.started":"2023-07-30T08:01:15.885996Z","shell.execute_reply":"2023-07-30T08:01:16.812973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.xlabel(\"Sentence Length\")\nplt.ylabel(\"Frequency\")\nplt.title(\"Sentence Length Distribuition\")\nplt.hist(x)","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:01:20.741299Z","iopub.execute_input":"2023-07-30T08:01:20.741815Z","iopub.status.idle":"2023-07-30T08:01:21.269524Z","shell.execute_reply.started":"2023-07-30T08:01:20.741761Z","shell.execute_reply":"2023-07-30T08:01:21.268034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the sentences have length<=100\n\nLet's look at the overall vocabulary size and the most frequent words\n","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\nvocab = {}\nfor sen in tqdm(df.sentence):\n    for j in sen.split(\" \"):\n        try:\n            vocab[j]+=1\n        except:\n            vocab[j]=1\nprint(\"Total words in vocabulary : \",len(vocab))","metadata":{"execution":{"iopub.status.busy":"2023-08-01T06:23:40.235972Z","iopub.execute_input":"2023-08-01T06:23:40.236401Z","iopub.status.idle":"2023-08-01T06:23:46.015997Z","shell.execute_reply.started":"2023-08-01T06:23:40.236367Z","shell.execute_reply":"2023-08-01T06:23:46.015059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So total 235761 words in vocabulary. That's huge. Let's look at the most frequent ones","metadata":{}},{"cell_type":"code","source":"sorted_vocab = sorted(vocab.items(),key = lambda kv:kv[1],reverse=True)\nsorted_vocab[:30]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T06:23:49.427311Z","iopub.execute_input":"2023-08-01T06:23:49.427940Z","iopub.status.idle":"2023-08-01T06:23:49.526867Z","shell.execute_reply.started":"2023-08-01T06:23:49.427907Z","shell.execute_reply":"2023-08-01T06:23:49.525496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So mostly bengali pronouns and common verbs.**","metadata":{}},{"cell_type":"markdown","source":"This is to remind you that this is the vocabulary of **train+validation**. You might wanna check how much **out of vocabulary** words does the validation set have. Let's check it.","metadata":{}},{"cell_type":"code","source":"def vocabulary(df):\n    vocab = {}\n    for sen in tqdm(df.sentence):\n        for j in sen.split(\" \"):\n            try:\n                vocab[j]+=1\n            except:\n                vocab[j]=1\n    return vocab","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:02:24.942247Z","iopub.execute_input":"2023-07-30T08:02:24.942729Z","iopub.status.idle":"2023-07-30T08:02:24.949622Z","shell.execute_reply.started":"2023-07-30T08:02:24.942690Z","shell.execute_reply":"2023-07-30T08:02:24.948331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = df[df[\"split\"]==\"train\"]\nval = df[df[\"split\"]==\"valid\"]\n\ntrain_vocab = vocabulary(train)\nval_vocab = vocabulary(val)\nprint(\"Total words in train vocabulary : \",len(train_vocab))\nprint(\"Total words in validation vocabulary : \",len(val_vocab))\n\ntrain_words = set([key for key,value in train_vocab.items()])\nvocab_words = set([key for key,value in val_vocab.items()])\n\nprint(\"Total Out of vocabulary words : \",len(vocab_words-train_words))\n","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:02:28.413950Z","iopub.execute_input":"2023-07-30T08:02:28.414514Z","iopub.status.idle":"2023-07-30T08:02:34.341676Z","shell.execute_reply.started":"2023-07-30T08:02:28.414473Z","shell.execute_reply":"2023-07-30T08:02:34.340241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Okay not so much words.\n\nNow let's look at the characters present.","metadata":{}},{"cell_type":"code","source":"chars = {}\nfor sen in tqdm(df.sentence):\n    for j in sen:\n        try:\n            chars[j]+=1\n        except:\n            chars[j]=1\nchars","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:23:31.969328Z","iopub.execute_input":"2023-08-01T07:23:31.969920Z","iopub.status.idle":"2023-08-01T07:23:50.067356Z","shell.execute_reply.started":"2023-08-01T07:23:31.969880Z","shell.execute_reply":"2023-08-01T07:23:50.065823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Total characters :\",len(chars))","metadata":{"execution":{"iopub.status.busy":"2023-07-30T08:04:21.096909Z","iopub.execute_input":"2023-07-30T08:04:21.098334Z","iopub.status.idle":"2023-07-30T08:04:21.104888Z","shell.execute_reply.started":"2023-07-30T08:04:21.098281Z","shell.execute_reply":"2023-07-30T08:04:21.103363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from __future__ import unicode_literals\n# import matplotlib.font_manager as font_manager\n# font = font_manager.FontProperties(fname=\"/kaggle/input/kalpurush-fonts/kalpurush-2.ttf\")\n\n# sns.distplot(x = \"x\", kde = True, data = sorted_chars)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:47:48.602516Z","iopub.execute_input":"2023-08-01T07:47:48.603010Z","iopub.status.idle":"2023-08-01T07:47:48.673562Z","shell.execute_reply.started":"2023-08-01T07:47:48.602973Z","shell.execute_reply":"2023-08-01T07:47:48.671856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.font_manager as font_manager\n\nprop = font_manager.FontProperties()\nfont = font_manager.FontProperties(fname=\"/kaggle/input/kalpurush-fonts/kalpurush-2.ttf\")\n\nsorted_chars = sorted(chars.items(),key = lambda x:x[1],reverse = True)\nprint(sorted_chars[:10])\nsorted_chars = dict(sorted_chars)\nplt.figure(figsize=(15,8))\nplt.title('Character Occurance')\nax = plt.bar(sorted_chars.keys(),sorted_chars.values())\nticklabels = plt.xticks()\nticks = []\nfor i in range(len(chars)):\n    ticks = ticklabels[1][i]._text\n# plt.xticks(\n#     ticks,\n#     fontproperties=font,\n# )\nplt.xlabel('Unique characters')\nplt.ylabel('Sample Count')\nplt.grid(axis='y')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:34:04.660104Z","iopub.execute_input":"2023-08-01T07:34:04.660648Z","iopub.status.idle":"2023-08-01T07:34:06.232537Z","shell.execute_reply.started":"2023-08-01T07:34:04.660609Z","shell.execute_reply":"2023-08-01T07:34:06.231369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Wordcloud","metadata":{}},{"cell_type":"code","source":"def Convert(tup):\n    d = {}\n    for a, b in tup:\n        d[a] = b\n    return d\n \nvocab_dict = Convert(sorted_vocab)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T06:25:32.117738Z","iopub.execute_input":"2023-08-01T06:25:32.118198Z","iopub.status.idle":"2023-08-01T06:25:32.267991Z","shell.execute_reply.started":"2023-08-01T06:25:32.118163Z","shell.execute_reply":"2023-08-01T06:25:32.266819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install git+https://github.com/csebuetnlp/normalizer","metadata":{"execution":{"iopub.status.busy":"2023-07-31T16:23:31.393418Z","iopub.execute_input":"2023-07-31T16:23:31.393760Z","iopub.status.idle":"2023-07-31T16:23:48.908544Z","shell.execute_reply.started":"2023-07-31T16:23:31.393732Z","shell.execute_reply":"2023-07-31T16:23:48.907156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from normalizer import normalize","metadata":{"execution":{"iopub.status.busy":"2023-07-31T16:23:48.911205Z","iopub.execute_input":"2023-07-31T16:23:48.911580Z","iopub.status.idle":"2023-07-31T16:23:49.065581Z","shell.execute_reply.started":"2023-07-31T16:23:48.911548Z","shell.execute_reply":"2023-07-31T16:23:49.064437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sentences = \"\"\ndef normalize_text(sent):\n    global sentences\n    normalized_sent = ''\n    for word in sent.split(' '):\n        normalized_sent += normalize(word) + \" \"\n    sentences += normalized_sent\n    return normalized_sent.rstrip(' ')","metadata":{"execution":{"iopub.status.busy":"2023-07-31T16:23:49.723686Z","iopub.execute_input":"2023-07-31T16:23:49.724050Z","iopub.status.idle":"2023-07-31T16:23:49.730364Z","shell.execute_reply.started":"2023-07-31T16:23:49.724014Z","shell.execute_reply":"2023-07-31T16:23:49.729270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from spacy.lang.bn import Bengali\nfrom spacy.lang.bn import STOP_WORDS as bengali_stopwords\nfrom collections import Counter","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:04:42.329477Z","iopub.execute_input":"2023-08-01T07:04:42.330479Z","iopub.status.idle":"2023-08-01T07:04:42.337521Z","shell.execute_reply.started":"2023-08-01T07:04:42.330436Z","shell.execute_reply":"2023-08-01T07:04:42.335880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\ntqdm.pandas()\n\ndf['normalized_sentence'] = df['sentence'].progress_apply(lambda x: normalize_text(x))","metadata":{"execution":{"iopub.status.busy":"2023-07-31T16:25:40.839471Z","iopub.execute_input":"2023-07-31T16:25:40.839874Z","iopub.status.idle":"2023-08-01T00:34:05.658875Z","shell.execute_reply.started":"2023-07-31T16:25:40.839818Z","shell.execute_reply":"2023-08-01T00:34:05.657058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bengali_nlp = Bengali()\nbengali_nlp.max_length(len(sentences)*2)\nbengali_doc = bengali_nlp(sentences)\nbengali_tokens = [token.text for token in bengali_doc]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T00:34:05.661055Z","iopub.execute_input":"2023-08-01T00:34:05.661417Z","iopub.status.idle":"2023-08-01T00:34:06.056581Z","shell.execute_reply.started":"2023-08-01T00:34:05.661391Z","shell.execute_reply":"2023-08-01T00:34:06.054618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bengali_tokens_counter = Counter(bengali_tokens)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T00:34:06.057885Z","iopub.status.idle":"2023-08-01T00:34:06.058310Z","shell.execute_reply.started":"2023-08-01T00:34:06.058126Z","shell.execute_reply":"2023-08-01T00:34:06.058147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:20:20.945708Z","iopub.execute_input":"2023-07-31T07:20:20.946081Z","iopub.status.idle":"2023-07-31T07:20:20.957978Z","shell.execute_reply.started":"2023-07-31T07:20:20.946047Z","shell.execute_reply":"2023-07-31T07:20:20.956872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install bnlp_toolkit","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:04:04.667675Z","iopub.execute_input":"2023-08-01T07:04:04.668094Z","iopub.status.idle":"2023-08-01T07:04:23.840608Z","shell.execute_reply.started":"2023-08-01T07:04:04.668064Z","shell.execute_reply":"2023-08-01T07:04:23.839221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from bnlp.corpus import stopwords, punctuations","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:04:23.843934Z","iopub.execute_input":"2023-08-01T07:04:23.844442Z","iopub.status.idle":"2023-08-01T07:04:25.427556Z","shell.execute_reply.started":"2023-08-01T07:04:23.844402Z","shell.execute_reply":"2023-08-01T07:04:25.426109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for word in bengali_stopwords:\n    try:\n        vocab_dict.pop(word)\n    except Exception as e:\n        print(e)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:11:00.657079Z","iopub.execute_input":"2023-08-01T07:11:00.657565Z","iopub.status.idle":"2023-08-01T07:11:00.666961Z","shell.execute_reply.started":"2023-08-01T07:11:00.657529Z","shell.execute_reply":"2023-08-01T07:11:00.665415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cmap = sns.color_palette(\"hls\", 8, as_cmap=True)\nwordcloud = WordCloud(font_path='/kaggle/input/kalpurush-fonts/kalpurush-2.ttf',\n                      width = 3000, height = 2000, \n                      max_words = 100, max_font_size= 1000,\n                      background_color = \"white\", colormap = cmap, \n                      collocations=False, \n                      stopwords=bengali_stopwords,\n                      random_state=36).generate_from_frequencies(vocab_dict)\nplt.figure( figsize=(10,5) )\nplt.imshow(wordcloud, interpolation='bilinear')\nplt.axis(\"off\")\nplt.margins(x=0, y=0)\nplt.show()\nplt.savefig('wordcloud.png')","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:14:21.621501Z","iopub.execute_input":"2023-08-01T07:14:21.622004Z","iopub.status.idle":"2023-08-01T07:14:30.349171Z","shell.execute_reply.started":"2023-08-01T07:14:21.621970Z","shell.execute_reply":"2023-08-01T07:14:30.347778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wordcloud.to_file(\"wordcloud1.png\")","metadata":{"execution":{"iopub.status.busy":"2023-08-01T07:21:50.350430Z","iopub.execute_input":"2023-08-01T07:21:50.351703Z","iopub.status.idle":"2023-08-01T07:21:52.394480Z","shell.execute_reply.started":"2023-08-01T07:21:50.351662Z","shell.execute_reply":"2023-08-01T07:21:52.393274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv('new_df.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-01T00:34:06.061241Z","iopub.status.idle":"2023-08-01T00:34:06.061565Z","shell.execute_reply.started":"2023-08-01T00:34:06.061425Z","shell.execute_reply":"2023-08-01T00:34:06.061439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n<div class=\"alert alert-block alert-info\" style=\"padding:25px; line-height: 1.7em; font-family: Verdana;\">\n    <center style=\"font-family: consolas; font-size: 32px; font-weight: bold;\"> Audio EDA</center>\n","metadata":{}},{"cell_type":"code","source":"!pip install mutagen","metadata":{"execution":{"iopub.status.busy":"2023-07-30T09:04:44.941750Z","iopub.execute_input":"2023-07-30T09:04:44.942317Z","iopub.status.idle":"2023-07-30T09:04:59.418438Z","shell.execute_reply.started":"2023-07-30T09:04:44.942275Z","shell.execute_reply":"2023-07-30T09:04:59.416490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport glob\nimport concurrent.futures\nfrom mutagen.mp3 import MP3\n\ndef process_mp3_file(mp3_file):\n    try:\n        audio = MP3(mp3_file)\n        duration = audio.info.length\n        sample_rate = audio.info.sample_rate\n        return mp3_file, duration, sample_rate\n    except Exception as e:\n        print(f\"Error processing {mp3_file}: {e}\")\n        return mp3_file, None, None\n\ndef process_mp3_files_in_batch(mp3_files):\n    durs = []\n    srs = []\n    with concurrent.futures.ThreadPoolExecutor() as executor:\n        futures = [executor.submit(process_mp3_file, mp3_file) for mp3_file in mp3_files]\n        for future in concurrent.futures.as_completed(futures):\n            result = future.result()\n            durs.append(result[1])\n            srs.append(result[2])\n    return durs,srs\n\ndef durations(mp3_files_directory):\n    # Get a list of all MP3 files in the directory\n    mp3_files = glob.glob(os.path.join(mp3_files_directory, \"*.mp3\"))\n\n    # Split the list of files into batches for processing\n    batch_size = 1000\n    file_batches = [mp3_files[i:i + batch_size] for i in tqdm(range(0, len(mp3_files), batch_size))]\n\n    durs = []\n    srs = []\n    for batch in tqdm(file_batches):\n        dur,sr = process_mp3_files_in_batch(batch)\n        durs.extend(dur)\n        srs.extend(sr)\n    return durs,srs","metadata":{"execution":{"iopub.status.busy":"2023-07-30T09:04:59.422459Z","iopub.execute_input":"2023-07-30T09:04:59.423401Z","iopub.status.idle":"2023-07-30T09:04:59.441246Z","shell.execute_reply.started":"2023-07-30T09:04:59.423355Z","shell.execute_reply":"2023-07-30T09:04:59.437703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_path = \"/kaggle/input/bengaliai-speech/train_mp3s\"\ndurs,srs = durations(train_path)","metadata":{"execution":{"iopub.status.busy":"2023-07-30T09:04:59.443389Z","iopub.execute_input":"2023-07-30T09:04:59.443871Z","iopub.status.idle":"2023-07-30T09:26:47.793515Z","shell.execute_reply.started":"2023-07-30T09:04:59.443834Z","shell.execute_reply":"2023-07-30T09:26:47.791664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.xlabel(\"Audio Length in seconds\")\nplt.ylabel(\"Frequency\")\nplt.title(\"Audio Length Distribuition\")\nplt.hist(durs)","metadata":{"execution":{"iopub.status.busy":"2023-07-30T09:26:47.796994Z","iopub.execute_input":"2023-07-30T09:26:47.797439Z","iopub.status.idle":"2023-07-30T09:26:53.600448Z","shell.execute_reply.started":"2023-07-30T09:26:47.797401Z","shell.execute_reply":"2023-07-30T09:26:53.597114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nice! Staircase.","metadata":{}},{"cell_type":"markdown","source":"For sampling rate :","metadata":{}},{"cell_type":"code","source":"set(srs)","metadata":{"execution":{"iopub.status.busy":"2023-07-30T09:26:53.602145Z","iopub.execute_input":"2023-07-30T09:26:53.602553Z","iopub.status.idle":"2023-07-30T09:26:53.624467Z","shell.execute_reply.started":"2023-07-30T09:26:53.602521Z","shell.execute_reply":"2023-07-30T09:26:53.622904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import librosa\nfrom pydub import AudioSegment\nimport numpy as np\n\n# audio_id = '000005f3362c'\naudio_id = df.iloc[np.random.randint(0, len(df))]['id']\n\npath = f\"/kaggle/input/bengaliai-speech/train_mp3s/{audio_id}.mp3\"\ndisplay(AudioSegment.from_file(path))\nprint(\"dataframe_id\", df[df['id'] == audio_id]['sentence'])\n\naudio , _ = librosa.load(path)\nprint(\"array:\", audio)\nprint(\"array length:\", len(audio))\nplt.figure(figsize = (100 , 30))\nsns.lineplot(audio)\n","metadata":{"execution":{"iopub.status.busy":"2023-08-07T08:15:54.569537Z","iopub.execute_input":"2023-08-07T08:15:54.570504Z","iopub.status.idle":"2023-08-07T08:15:58.576003Z","shell.execute_reply.started":"2023-08-07T08:15:54.570464Z","shell.execute_reply":"2023-08-07T08:15:58.575103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## fast fourier transformation to get frequency information \nN = len(audio)\nw = np.hamming(N)\nX = np.fft.fft(audio*N)\nsns.lineplot(X)\n# Sending abs(X) to avoid imaginary number\nsns.lineplot(np.abs(X))","metadata":{"execution":{"iopub.status.busy":"2023-08-07T08:16:03.169283Z","iopub.execute_input":"2023-08-07T08:16:03.170119Z","iopub.status.idle":"2023-08-07T08:16:04.117536Z","shell.execute_reply.started":"2023-08-07T08:16:03.170073Z","shell.execute_reply":"2023-08-07T08:16:04.116193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# taking the logarithm value as fine grain audio is getting lost\nsns.lineplot(np.log(np.abs(X)))","metadata":{"execution":{"iopub.status.busy":"2023-08-07T08:16:12.750705Z","iopub.execute_input":"2023-08-07T08:16:12.751161Z","iopub.status.idle":"2023-08-07T08:16:13.415683Z","shell.execute_reply.started":"2023-08-07T08:16:12.751124Z","shell.execute_reply":"2023-08-07T08:16:13.414832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}