{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"> The competition dataset might contain the audios of the [mozilla-foundation/common_voice_11_0](https://huggingface.co/datasets/mozilla-foundation/common_voice_11_0) dataset. In this notebook we'll try to figure out whether it's true.\n","metadata":{}},{"cell_type":"markdown","source":"# Load CommonVoice dataset","metadata":{}},{"cell_type":"markdown","source":"> Let's load the common voice dataset first from Huggingface","metadata":{}},{"cell_type":"code","source":"from datasets import load_dataset","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:15:34.924909Z","iopub.execute_input":"2023-07-25T19:15:34.925939Z","iopub.status.idle":"2023-07-25T19:15:36.151270Z","shell.execute_reply.started":"2023-07-25T19:15:34.925893Z","shell.execute_reply":"2023-07-25T19:15:36.149957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = load_dataset(\"mozilla-foundation/common_voice_11_0\", \"bn\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:15:36.153320Z","iopub.execute_input":"2023-07-25T19:15:36.154224Z","iopub.status.idle":"2023-07-25T19:26:13.311281Z","shell.execute_reply.started":"2023-07-25T19:15:36.154179Z","shell.execute_reply":"2023-07-25T19:26:13.310276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Let's look at the dataset","metadata":{}},{"cell_type":"code","source":"dataset","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:34:25.296701Z","iopub.execute_input":"2023-07-25T19:34:25.297284Z","iopub.status.idle":"2023-07-25T19:34:25.306611Z","shell.execute_reply.started":"2023-07-25T19:34:25.297248Z","shell.execute_reply":"2023-07-25T19:34:25.305064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":">                   It has 5 folds\n\n| **Train**     | Total audios |\n| ----------- | ----------- |\n| train      | 16777       |\n| test   | 8353        |\n| validation | 8353 |\n| other | 225826 |\n| invalidated | 6447 |","metadata":{}},{"cell_type":"markdown","source":"> Each sample has these columns\n``` 'client_id', 'path', 'audio', 'sentence', 'up_votes', 'down_votes', 'age', 'gender', 'accent', 'locale', 'segment' ```","metadata":{}},{"cell_type":"markdown","source":"**We are only interested in the sentences. We'll check which sentences among these are present in our dataset.**","metadata":{}},{"cell_type":"markdown","source":"Let's load the competition dataset first","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(\"/kaggle/input/bengaliai-speech/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:36:34.385650Z","iopub.execute_input":"2023-07-25T19:36:34.386347Z","iopub.status.idle":"2023-07-25T19:36:41.703950Z","shell.execute_reply.started":"2023-07-25T19:36:34.386306Z","shell.execute_reply":"2023-07-25T19:36:41.701728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = df[df[\"split\"]==\"train\"]\nval = df[df[\"split\"]==\"val\"]","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:36:41.707216Z","iopub.execute_input":"2023-07-25T19:36:41.708027Z","iopub.status.idle":"2023-07-25T19:36:42.215187Z","shell.execute_reply.started":"2023-07-25T19:36:41.707968Z","shell.execute_reply":"2023-07-25T19:36:42.213092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sens = train.sentence.tolist()\nval_sens = val.sentence.tolist()","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:36:42.217747Z","iopub.execute_input":"2023-07-25T19:36:42.218292Z","iopub.status.idle":"2023-07-25T19:36:42.268748Z","shell.execute_reply.started":"2023-07-25T19:36:42.218259Z","shell.execute_reply":"2023-07-25T19:36:42.267509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's make a dictionary of the sentences as Hashing is faster. ","metadata":{}},{"cell_type":"code","source":"def make_dict(sens):\n    dic = {}\n    for i in sens:\n        try:\n            dic[i]+=1\n        except:\n            dic[i]=1\n    return dic","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:37:21.874062Z","iopub.execute_input":"2023-07-25T19:37:21.874855Z","iopub.status.idle":"2023-07-25T19:37:21.885297Z","shell.execute_reply.started":"2023-07-25T19:37:21.874770Z","shell.execute_reply":"2023-07-25T19:37:21.883261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dict = make_dict(train_sens)\nval_dict = make_dict(val_sens)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:37:24.612051Z","iopub.execute_input":"2023-07-25T19:37:24.612592Z","iopub.status.idle":"2023-07-25T19:37:25.614978Z","shell.execute_reply.started":"2023-07-25T19:37:24.612549Z","shell.execute_reply":"2023-07-25T19:37:25.613582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"splits = [\"train\",\"test\",\"validation\",\"other\",\"invalidated\"]","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:37:57.073789Z","iopub.execute_input":"2023-07-25T19:37:57.074583Z","iopub.status.idle":"2023-07-25T19:37:57.083520Z","shell.execute_reply.started":"2023-07-25T19:37:57.074539Z","shell.execute_reply":"2023-07-25T19:37:57.081469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\ngrand_cnt = 0\ncv_total = 0\nfor split in tqdm(splits):\n    cnt_train = 0\n    cnt_val = 0\n    for sentence in dataset[split]['sentence']:\n        try:\n            train_dict[sentence]\n            cnt_train+=1\n        except:\n            continue\n        \n        try:\n            val_dict[sentence]\n            cnt_val+=1\n        except:\n            continue\n    \n    grand_cnt+=cnt_train+cnt_val\n    print(\"Split Name : \",split)\n    cv_total+= len(dataset[split]['sentence'])\n    print(f\"Total audios in commonvoice {split}: \",len(dataset[split]['sentence']))\n    print(\"Total audios in train : \",cnt_train)\n    print(\"Total audios in val : \",cnt_val)\n    print(\"-\"*80)\n            \nprint(f\"Total common voice audio :{cv_total}\\n Audios present here : {grand_cnt}\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:37:59.179460Z","iopub.execute_input":"2023-07-25T19:37:59.181613Z","iopub.status.idle":"2023-07-25T19:38:04.459703Z","shell.execute_reply.started":"2023-07-25T19:37:59.181517Z","shell.execute_reply":"2023-07-25T19:38:04.457313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Finding audios with overlapping sentences","metadata":{}},{"cell_type":"code","source":"cv_voc = {}\nfor split in tqdm(splits):\n    for sentence in dataset[split]['sentence']:\n        try:\n            cv_voc[sentence]+=1\n        except:\n            cv_voc[sentence]=1","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:50:32.429108Z","iopub.execute_input":"2023-07-25T19:50:32.429625Z","iopub.status.idle":"2023-07-25T19:50:33.347220Z","shell.execute_reply.started":"2023-07-25T19:50:32.429589Z","shell.execute_reply":"2023-07-25T19:50:33.345945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cnt = 0\nidxs = []\nfor idx,i in tqdm(enumerate(df.sentence.tolist())):\n    try:\n        cv_voc[i]\n        cnt+=1\n        idxs.append(idx)\n    except:\n        continue\ncnt\n        ","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:50:58.023809Z","iopub.execute_input":"2023-07-25T19:50:58.024351Z","iopub.status.idle":"2023-07-25T19:50:59.414685Z","shell.execute_reply.started":"2023-07-25T19:50:58.024313Z","shell.execute_reply":"2023-07-25T19:50:59.413357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"idx_df = pd.DataFrame({\"id\":idxs})\nidx_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:51:02.496870Z","iopub.execute_input":"2023-07-25T19:51:02.497342Z","iopub.status.idle":"2023-07-25T19:51:02.656155Z","shell.execute_reply.started":"2023-07-25T19:51:02.497309Z","shell.execute_reply":"2023-07-25T19:51:02.654656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"idx_df.to_csv(\"indexes.csv\",index=False)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T19:51:04.827224Z","iopub.execute_input":"2023-07-25T19:51:04.827688Z","iopub.status.idle":"2023-07-25T19:51:05.429875Z","shell.execute_reply.started":"2023-07-25T19:51:04.827653Z","shell.execute_reply":"2023-07-25T19:51:05.428852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are the indexes that contain sentences that overlap in both of the datasets!","metadata":{}},{"cell_type":"markdown","source":"# Hear some of the overlapping samples to check whether they actually match","metadata":{}},{"cell_type":"code","source":"import soundfile as sf\nfrom pydub import AudioSegment\nimport IPython.display as ipd","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cnt = 0\nbase_path = \"/kaggle/input/bengaliai-speech/train_mp3s/\"\n\nfor data in dataset[\"train\"]:\n    if data[\"sentence\"] in df.sentence.tolist():\n        audio = data[\"audio\"][\"array\"]\n        sentence = data[\"sentence\"]\n        sr = data[\"audio\"][\"sampling_rate\"]\n        idx = df[df[\"sentence\"]==sentence]\n        \n        if len(idx)>1:\n            continue\n        #idx = idx[idx[\"split\"]==\"train\"]\n        \n        print(\"Sentence : \",sentence)\n        print(\"Common Voice audio :\")\n        display(ipd.Audio(audio,rate=sr))\n        \n        if len(idx)>1:\n            print(\"Multiple audios in the competition dataset with the same sentence \\n\")\n        \n        for i in range(len(idx)):\n            path = base_path+idx['id'].iloc[i]+\".mp3\"\n            print(\"Competition data audio : \",idx['id'].iloc[i]+\".mp3\")\n            display(AudioSegment.from_file(path))\n            \n        print(\"-\"*80)\n        print(\"-\"*80)\n        \n        cnt+=1\n        \n        if cnt>=5:\n            break\n    ","metadata":{"execution":{"iopub.status.busy":"2023-07-25T20:11:17.004912Z","iopub.execute_input":"2023-07-25T20:11:17.005566Z","iopub.status.idle":"2023-07-25T20:11:22.409025Z","shell.execute_reply.started":"2023-07-25T20:11:17.005525Z","shell.execute_reply":"2023-07-25T20:11:22.407560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Well, there's a problem though. Some of the audios are the same audios, some aren't.\n\n> Some sentences have multiple audio samples. Let's look at some of them","metadata":{}},{"cell_type":"code","source":"cnt = 0\nbase_path = \"/kaggle/input/bengaliai-speech/train_mp3s/\"\n\nfor data in dataset[\"train\"]:\n    if data[\"sentence\"] in df.sentence.tolist():\n        audio = data[\"audio\"][\"array\"]\n        sentence = data[\"sentence\"]\n        sr = data[\"audio\"][\"sampling_rate\"]\n        idx = df[df[\"sentence\"]==sentence]\n        \n        if len(idx)!=3:\n            continue\n        #idx = idx[idx[\"split\"]==\"train\"]\n        \n        print(\"Sentence : \",sentence)\n        print(\"Common Voice audio :\")\n        display(ipd.Audio(audio,rate=sr))\n        \n        if len(idx)>1:\n            print(\"Multiple audios in the competition dataset with the same sentence \\n\")\n        \n        for i in range(len(idx)):\n            path = base_path+idx['id'].iloc[i]+\".mp3\"\n            print(\"Competition data audio : \",idx['id'].iloc[i]+\".mp3\")\n            display(AudioSegment.from_file(path))\n            \n        print(\"-\"*80)\n        print(\"-\"*80)\n        \n        cnt+=1\n        \n        if cnt>=5:\n            break\n    ","metadata":{"execution":{"iopub.status.busy":"2023-07-25T20:13:23.376689Z","iopub.execute_input":"2023-07-25T20:13:23.377229Z","iopub.status.idle":"2023-07-25T20:13:50.528402Z","shell.execute_reply.started":"2023-07-25T20:13:23.377189Z","shell.execute_reply":"2023-07-25T20:13:50.527019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well here for each case, each of the examples of our dataset has three audio samples per sentence. Of them one was present in the Common Voice dataset.","metadata":{}},{"cell_type":"markdown","source":"# Let's look at an example in details","metadata":{}},{"cell_type":"code","source":"tmp = df[df[\"sentence\"]=='এই ভোগ সম্পূর্ণভাবে নির্ভর করে তার নিজের উপরে।']\ntmp","metadata":{"execution":{"iopub.status.busy":"2023-07-25T20:08:38.524629Z","iopub.execute_input":"2023-07-25T20:08:38.526927Z","iopub.status.idle":"2023-07-25T20:08:38.805252Z","shell.execute_reply.started":"2023-07-25T20:08:38.526860Z","shell.execute_reply":"2023-07-25T20:08:38.802911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> So this sentence has two occurence in the competition dataset. One is present in validation and the other one is in the train set. Let's see which one is the same to the CV dataset.","metadata":{}},{"cell_type":"code","source":"for data in dataset[\"train\"]:\n    if data['sentence']== 'এই ভোগ সম্পূর্ণভাবে নির্ভর করে তার নিজের উপরে।':\n        x = data['audio']['array']\n        break","metadata":{"execution":{"iopub.status.busy":"2023-07-25T20:06:26.383783Z","iopub.execute_input":"2023-07-25T20:06:26.384334Z","iopub.status.idle":"2023-07-25T20:06:26.401350Z","shell.execute_reply.started":"2023-07-25T20:06:26.384291Z","shell.execute_reply":"2023-07-25T20:06:26.399927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Audio in Common Voice dataset : \")\nipd.Audio(x,rate=48000)","metadata":{"execution":{"iopub.status.busy":"2023-07-25T20:07:01.392276Z","iopub.execute_input":"2023-07-25T20:07:01.392766Z","iopub.status.idle":"2023-07-25T20:07:01.414031Z","shell.execute_reply.started":"2023-07-25T20:07:01.392731Z","shell.execute_reply":"2023-07-25T20:07:01.412297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Let's now hear the two audios in our dataset**","metadata":{}},{"cell_type":"code","source":"file_id = tmp[tmp[\"split\"]==\"train\"][\"id\"].tolist()[0]\nAudioSegment.from_file(base_path+file_id+\".mp3\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T20:09:33.802185Z","iopub.execute_input":"2023-07-25T20:09:33.802823Z","iopub.status.idle":"2023-07-25T20:09:34.176385Z","shell.execute_reply.started":"2023-07-25T20:09:33.802662Z","shell.execute_reply":"2023-07-25T20:09:34.174788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Well, this one doesn't match with the previous audio. Let's look at the other one**","metadata":{}},{"cell_type":"code","source":"file_id = tmp[tmp[\"split\"]==\"valid\"][\"id\"].tolist()[0]\nAudioSegment.from_file(base_path+file_id+\".mp3\")","metadata":{"execution":{"iopub.status.busy":"2023-07-25T20:10:16.634625Z","iopub.execute_input":"2023-07-25T20:10:16.635115Z","iopub.status.idle":"2023-07-25T20:10:16.958889Z","shell.execute_reply.started":"2023-07-25T20:10:16.635078Z","shell.execute_reply":"2023-07-25T20:10:16.957865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Now this one is the exact match with the common voice sample!","metadata":{}}]}