{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"``` \nThe release of the train metadata will have an heavy impact on the training performance. Since the dataset is MaCro data, it might contain wrong annotations or biased annotations. Having these audios in the training might confuse the models. This metadata might help us filter out those data from the training set. Let's explore the metadata today and see how much of data are actually helpful\n```","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport soundfile as sf\nfrom pydub import AudioSegment\nfrom tqdm.notebook import tqdm\ntqdm.pandas()\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:45:53.558525Z","iopub.execute_input":"2023-08-30T19:45:53.559941Z","iopub.status.idle":"2023-08-30T19:45:53.567314Z","shell.execute_reply.started":"2023-08-30T19:45:53.559892Z","shell.execute_reply":"2023-08-30T19:45:53.566363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata_path = \"/kaggle/input/bengaliai-speech-train-metadata/train_metadata.csv\"\ndf_meta = pd.read_csv(metadata_path)\nprint(\"metadata shape :\",df_meta.shape)\nprint(\"Columns in metadata : \",df_meta.columns)\ndisplay(df_meta.head())","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:45:53.57297Z","iopub.execute_input":"2023-08-30T19:45:53.573686Z","iopub.status.idle":"2023-08-30T19:46:14.744131Z","shell.execute_reply.started":"2023-08-30T19:45:53.573648Z","shell.execute_reply":"2023-08-30T19:46:14.743025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well there are a lot of features here, focusing on the quality of the audios and the transcriptions! We'll mostly focus on the quality of the transcriptions here to find out the consistent audio and ground truth transcriptions!","metadata":{}},{"cell_type":"code","source":"all_cols = df_meta.columns\nall_cols","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:46:14.746283Z","iopub.execute_input":"2023-08-30T19:46:14.746988Z","iopub.status.idle":"2023-08-30T19:46:14.754215Z","shell.execute_reply.started":"2023-08-30T19:46:14.746949Z","shell.execute_reply":"2023-08-30T19:46:14.753024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns_to_take = ['id','ggl_cer','ggl_wer','google_preds','sentence','yellowking_preds','ykg_wer','ykg_cer']\nmode = 'not_debug'\nif mode=='debug':\n    df = df_meta.iloc[:1000]\nelse:\n    df = df_meta","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:46:14.755395Z","iopub.execute_input":"2023-08-30T19:46:14.755722Z","iopub.status.idle":"2023-08-30T19:46:15.504423Z","shell.execute_reply.started":"2023-08-30T19:46:14.755695Z","shell.execute_reply":"2023-08-30T19:46:15.503201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One important thing here is yellowking_model's outputs are unicode normalized, whereas the ground truths aren't. so to compare the yellowking prediction with the ground truth, we need have both of them in the same canonical form. So let's normalize all of them first.","metadata":{}},{"cell_type":"code","source":"%%capture tst\n!pip install bnunicodenormalizer\nfrom bnunicodenormalizer import Normalizer ","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:46:15.507614Z","iopub.execute_input":"2023-08-30T19:46:15.508162Z","iopub.status.idle":"2023-08-30T19:46:27.990937Z","shell.execute_reply.started":"2023-08-30T19:46:15.508122Z","shell.execute_reply":"2023-08-30T19:46:27.989378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bnorm = Normalizer()\ndef normalize(sen):\n    \n    try:\n        _words = [bnorm(word)['normalized']  for word in sen.split()]\n        return \" \".join([word for word in _words if word is not None])\n    except:\n        sen = \"!\"\n        return sen","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:47:45.345753Z","iopub.execute_input":"2023-08-30T19:47:45.346365Z","iopub.status.idle":"2023-08-30T19:47:45.353488Z","shell.execute_reply.started":"2023-08-30T19:47:45.346292Z","shell.execute_reply":"2023-08-30T19:47:45.352452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['yellowking_preds'] = df['yellowking_preds'].progress_apply(lambda x:normalize(x))\ndf['sentence'] = df['sentence'].progress_apply(lambda x:normalize(x))\ndf['google_preds'] = df['google_preds'].progress_apply(lambda x:normalize(x))\n","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:47:45.361314Z","iopub.execute_input":"2023-08-30T19:47:45.362002Z","iopub.status.idle":"2023-08-30T19:47:56.839778Z","shell.execute_reply.started":"2023-08-30T19:47:45.361964Z","shell.execute_reply":"2023-08-30T19:47:56.837851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture ts\n!pip install jiwer\nfrom jiwer import wer, cer","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:47:56.842004Z","iopub.execute_input":"2023-08-30T19:47:56.842643Z","iopub.status.idle":"2023-08-30T19:48:10.093626Z","shell.execute_reply.started":"2023-08-30T19:47:56.842604Z","shell.execute_reply":"2023-08-30T19:48:10.092256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's calculate the new WER between yellowking model and ground truth sentence","metadata":{}},{"cell_type":"code","source":"yk_wer = []\nfor i in tqdm(range(len(df))):\n    yk_wer.append(wer(df['sentence'].iloc[i],df['yellowking_preds'].iloc[i]))\n    \ndf['yk_wer_new'] = yk_wer\nsum(df['yk_wer_new'] == df['ykg_wer'])","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:59:20.603795Z","iopub.execute_input":"2023-08-30T19:59:20.60426Z","iopub.status.idle":"2023-08-30T19:59:20.788626Z","shell.execute_reply.started":"2023-08-30T19:59:20.604223Z","shell.execute_reply":"2023-08-30T19:59:20.787582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"faulty = df[df['yk_wer_new']!=df['ykg_wer']]\nfaulty[['sentence','yellowking_preds','ykg_wer','yk_wer_new']].head()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T19:59:23.164792Z","iopub.execute_input":"2023-08-30T19:59:23.165251Z","iopub.status.idle":"2023-08-30T19:59:23.183703Z","shell.execute_reply.started":"2023-08-30T19:59:23.165213Z","shell.execute_reply":"2023-08-30T19:59:23.182716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":">  Well , we can see that the train metadata provided and the the wer aren't same here.","metadata":{}},{"cell_type":"code","source":"ggl_wer = []\nfor i in tqdm(range(len(df))):\n    ggl_wer.append(wer(df['sentence'].iloc[i],df['google_preds'].iloc[i]))\n    \ndf['ggl_wer_new'] = ggl_wer\nsum(df['ggl_wer_new'] == df['ggl_wer'])","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:02:30.571376Z","iopub.execute_input":"2023-08-30T20:02:30.571877Z","iopub.status.idle":"2023-08-30T20:02:30.785006Z","shell.execute_reply.started":"2023-08-30T20:02:30.57183Z","shell.execute_reply":"2023-08-30T20:02:30.784257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Let's replace the ykg_wer and ggl_wer column with the new values","metadata":{}},{"cell_type":"code","source":"df['ykg_wer'] = df['yk_wer_new']\ndf['ggl_wer'] = df['ggl_wer_new']","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:05:55.815883Z","iopub.execute_input":"2023-08-30T20:05:55.816951Z","iopub.status.idle":"2023-08-30T20:05:55.823166Z","shell.execute_reply.started":"2023-08-30T20:05:55.816904Z","shell.execute_reply":"2023-08-30T20:05:55.822229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df[all_cols]\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:05:55.824689Z","iopub.execute_input":"2023-08-30T20:05:55.824964Z","iopub.status.idle":"2023-08-30T20:05:55.851468Z","shell.execute_reply.started":"2023-08-30T20:05:55.824941Z","shell.execute_reply":"2023-08-30T20:05:55.850382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:06:02.569042Z","iopub.execute_input":"2023-08-30T20:06:02.569988Z","iopub.status.idle":"2023-08-30T20:06:02.576369Z","shell.execute_reply.started":"2023-08-30T20:06:02.56995Z","shell.execute_reply":"2023-08-30T20:06:02.575391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv(\"train_metadata_corrected.csv\",index=False)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T20:05:55.852708Z","iopub.execute_input":"2023-08-30T20:05:55.85316Z","iopub.status.idle":"2023-08-30T20:05:55.891106Z","shell.execute_reply.started":"2023-08-30T20:05:55.853113Z","shell.execute_reply":"2023-08-30T20:05:55.890157Z"},"trusted":true},"execution_count":null,"outputs":[]}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}