{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Bengali Example Audio Wav2Vec2\nhttps://www.kaggle.com/code/stpeteishii/english-audio-wav2vec2<br/>\nhttps://www.kaggle.com/stpeteishii/bengali-example-audio-indicwav2vec","metadata":{"papermill":{"duration":0.009089,"end_time":"2023-08-21T04:35:41.194748","exception":false,"start_time":"2023-08-21T04:35:41.185659","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## The YellowKing model \nThe YellowKing model is a speech recognition model that was developed by a team of researchers from the Indian Institute of Technology, Madras. It was the champion on the Kaggle community competition \"DL Sprint\" in 2022, and achieved a Levenshtein Distance of 6.2349 on a hidden dataset, which was 1.1049 units lower than other competing submissions.\nThe YellowKing model is based on the wav2vec2 architecture, which is a neural network that is specifically designed for speech recognition. The model was trained on a large dataset of Bengali speech recordings, and it is able to recognize a variety of Bengali words and phrases.\nThe YellowKing model is a promising new development in speech recognition technology. It is able to achieve high accuracy on a variety of tasks, and it is still under development, so it is likely to improve even further in the future.\n","metadata":{"papermill":{"duration":0.007699,"end_time":"2023-08-21T04:35:41.211977","exception":false,"start_time":"2023-08-21T04:35:41.204278","status":"completed"},"tags":[]}},{"cell_type":"code","source":"!pip install moviepy","metadata":{"papermill":{"duration":12.378265,"end_time":"2023-08-21T04:35:53.598356","exception":false,"start_time":"2023-08-21T04:35:41.220091","status":"completed"},"tags":[],"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-08-21T08:10:01.375365Z","iopub.execute_input":"2023-08-21T08:10:01.375836Z","iopub.status.idle":"2023-08-21T08:10:08.221346Z","shell.execute_reply.started":"2023-08-21T08:10:01.375798Z","shell.execute_reply":"2023-08-21T08:10:08.220086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from transformers import Wav2Vec2Tokenizer, Wav2Vec2ForCTC\nfrom IPython.display import Audio\nfrom IPython.display import Video\nimport moviepy.editor as mp\nimport torch\nimport librosa\nimport os\nimport pandas as pd","metadata":{"papermill":{"duration":4.683807,"end_time":"2023-08-21T04:35:58.306083","exception":false,"start_time":"2023-08-21T04:35:53.622276","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-08-21T08:10:08.223446Z","iopub.execute_input":"2023-08-21T08:10:08.223770Z","iopub.status.idle":"2023-08-21T08:10:08.229525Z","shell.execute_reply.started":"2023-08-21T08:10:08.223739Z","shell.execute_reply":"2023-08-21T08:10:08.228329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenizer = Wav2Vec2Tokenizer.from_pretrained(\"/kaggle/input/yellowking-dlsprint-model/YellowKing_processor\")\nmodel = Wav2Vec2ForCTC.from_pretrained(\"/kaggle/input/yellowking-dlsprint-model/YellowKing_model\")","metadata":{"execution":{"iopub.status.busy":"2023-08-21T08:10:08.231434Z","iopub.execute_input":"2023-08-21T08:10:08.231930Z","iopub.status.idle":"2023-08-21T08:10:08.251626Z","shell.execute_reply.started":"2023-08-21T08:10:08.231881Z","shell.execute_reply":"2023-08-21T08:10:08.250592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clip_paths=[]\nfor dirname, _, filenames in os.walk('/kaggle/input/bengaliai-speech/examples'):\n    for filename in filenames:\n        clip_paths+=[(os.path.join(dirname, filename))]","metadata":{"papermill":{"duration":0.056286,"end_time":"2023-08-21T04:36:37.050482","exception":false,"start_time":"2023-08-21T04:36:36.994196","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-08-21T08:10:19.499282Z","iopub.execute_input":"2023-08-21T08:10:19.499756Z","iopub.status.idle":"2023-08-21T08:10:19.509187Z","shell.execute_reply.started":"2023-08-21T08:10:19.499706Z","shell.execute_reply":"2023-08-21T08:10:19.505404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#cc = []\nexdata=pd.DataFrame(columns=['path','sentence'],index=range(len(clip_paths)))\nfor i,path in enumerate(clip_paths):\n    print(path.split('/')[-1])\n    # Load the audio with the librosa library\n    input_audio, _ = librosa.load(path, sr=16000)\n    # Tokenize the audio\n    input_values = tokenizer(input_audio, return_tensors=\"pt\", padding=\"longest\").input_values\n    # Feed it through Wav2Vec & choose the most probable tokens\n    with torch.no_grad():\n        logits = model(input_values).logits\n        predicted_ids = torch.argmax(logits, dim=-1)\n    # Decode & add to our caption string\n    transcription = tokenizer.batch_decode(predicted_ids)[0]\n    #cc += [transcription]\n    exdata.loc[i,'sentence']=transcription\n    exdata.loc[i,'path']=path","metadata":{"papermill":{"duration":875.02182,"end_time":"2023-08-21T04:51:12.097553","exception":false,"start_time":"2023-08-21T04:36:37.075733","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-08-21T08:10:19.511404Z","iopub.execute_input":"2023-08-21T08:10:19.511849Z","iopub.status.idle":"2023-08-21T08:19:38.682889Z","shell.execute_reply.started":"2023-08-21T08:10:19.511801Z","shell.execute_reply":"2023-08-21T08:19:38.681813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Audio(clip_paths[0])","metadata":{"papermill":{"duration":0.367026,"end_time":"2023-08-21T04:51:12.494135","exception":false,"start_time":"2023-08-21T04:51:12.127109","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-08-21T08:19:38.684653Z","iopub.execute_input":"2023-08-21T08:19:38.684979Z","iopub.status.idle":"2023-08-21T08:19:38.826825Z","shell.execute_reply.started":"2023-08-21T08:19:38.684945Z","shell.execute_reply":"2023-08-21T08:19:38.825550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(exdata)","metadata":{"papermill":{"duration":0.31888,"end_time":"2023-08-21T04:51:13.122786","exception":false,"start_time":"2023-08-21T04:51:12.803906","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-08-21T08:19:38.828715Z","iopub.execute_input":"2023-08-21T08:19:38.829052Z","iopub.status.idle":"2023-08-21T08:19:38.841978Z","shell.execute_reply.started":"2023-08-21T08:19:38.829018Z","shell.execute_reply":"2023-08-21T08:19:38.841069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This result seems better than that of indicwav2vec.<br/>\nhttps://www.kaggle.com/stpeteishii/bengali-example-audio-indicwav2vec","metadata":{}},{"cell_type":"code","source":"exdata.to_csv('exdata.csv',index=False)","metadata":{"papermill":{"duration":0.286266,"end_time":"2023-08-21T04:51:13.680765","exception":false,"start_time":"2023-08-21T04:51:13.394499","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-08-21T08:19:38.843454Z","iopub.execute_input":"2023-08-21T08:19:38.843995Z","iopub.status.idle":"2023-08-21T08:19:38.854873Z","shell.execute_reply.started":"2023-08-21T08:19:38.843955Z","shell.execute_reply":"2023-08-21T08:19:38.853983Z"},"trusted":true},"execution_count":null,"outputs":[]}]}