{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#006600; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #003300\">Āmi tōmākē bhālōbāsi</p>\n\n<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n    \n`Āmi tōmākē bhālōbāsi` means `I Love You` in $Bengali$\n    \nOne interesting fact about Bengali Language. **Rabindranath Tagore, who wrote Jan Gan Man, National Antehm of India, was a Bengali Writer**\n    \n$Bengali$ is an $Indo-Aryan$ language spoken by over $220$ $Million$ people as a `first`/`second` language. It is the `official language` of $Bangladesh$ and one of the $22$ `Scheduled Languages` of $India$(my country $:)$). $Bengali$ is the `most widely` spoken language in $Bangladesh$, with over $100$ $Million$ speakers. It is also the `second-most widely` spoken `language` in $India$, after $Hindi$. $Bengali$ is written in the $Bengali$ $Script$, which is a $Brahmic$ $Script$. The $Bengali$ $Script$ is derived from the $Devanagari$ $Script$, but it has some `unique features`. $Bengali$ is a `tonal language`, which means that the `pitch of the voice` can `change the meaning of a word`. For example, the word `amar` can mean `my` or `not mine` depending on the `pitch of the voice`. $Bengali$ is a `rich` and `expressive language`. It has a `large vocabulary` and a `complex grammar`. $Bengali$ is also a very `musical language`, and it is often used in poetry and song.","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#ffff00 ; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #ffff00 \">1 | Goal ⚽️</p>\n\n<div style=\"border-radius:10px; border:#ffff00  solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n    \nAs it is written in the **[overview]()**, the error rate of $Google$ $Translator$ for $Bengali$ $Language$ is around $74$%, which is really huge. $Bengali.AI$, is hosting this competition to find better soltutions to the `translating system`, which will help in the communication between locals and tourists at many places.","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#800080; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #800080\">2 | Advisory 📃</p>\n\n<div style=\"border-radius:10px; border:#800080 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n    \n* The training data is huge, try to work on a smaller sample, to save resource and time. Once the pipeline is in the good condition, then send the whole data \n    \n* Padd the sentences to the maximum length, to avoid any shape errors\n    \n* The total amount of files in the `train_mp3s` were found to be $9,63,636$ which takes around $3$ Hours to load with `Librosa`. I have decided to go with $5,354$ samples which is $0.005$% of the Original Data. This took me around $15$ Seconds to load.\n* The length of all the audio files is the same which is $2,250$, denoting $2,250$, Seconds(I think so )","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#00FFFF; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #00FFFF\">3 | Data 💡</p>\n\n<div style=\"border-radius:10px; border:#00FFFF solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n    \nLets dive into the data ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns \nimport tqdm\nimport matplotlib.pyplot as plt \nimport librosa\nimport os","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-07-18T03:53:07.159638Z","iopub.execute_input":"2023-07-18T03:53:07.159994Z","iopub.status.idle":"2023-07-18T03:53:07.164286Z","shell.execute_reply.started":"2023-07-18T03:53:07.159968Z","shell.execute_reply":"2023-07-18T03:53:07.163468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#00FFFF solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\n## $3.1$ $|$ $Train$\nOur main training data is in the `train.csv` which is a `DataFrame` type object that contains $3$ columns\n* $ID$ - This is the unique ID given to each sample. In many cases ID becomes useless, but here it is the connecting point between the training point and the target data\n* $Sentences$ - This is the column that contains target data, which is in the sentence format\n* $Train/Valid$ - The last column shows, which part of the sample is training data and which is the validation data. Though we can make our own splits, but I think the host has provided this information, because they had better splits, over the distribution of the data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/bengaliai-speech/train.csv\")\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T03:18:53.216785Z","iopub.execute_input":"2023-07-18T03:18:53.217105Z","iopub.status.idle":"2023-07-18T03:18:56.026421Z","shell.execute_reply.started":"2023-07-18T03:18:53.217081Z","shell.execute_reply":"2023-07-18T03:18:56.025521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#00FFFF solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\n## $3.2$ $|$ $Train_-mp3s$\n    \nThis is a folder that contains our training samples. These are in audio files, and we wil thus use `Librosa` to read these files into array format","metadata":{}},{"cell_type":"code","source":"audio , _ = librosa.load('/kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3')\nprint(audio)\nplt.figure(figsize = (100 , 30))\nsns.lineplot(audio)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T03:56:13.424389Z","iopub.execute_input":"2023-07-18T03:56:13.424750Z","iopub.status.idle":"2023-07-18T03:56:15.632086Z","shell.execute_reply.started":"2023-07-18T03:56:13.424721Z","shell.execute_reply":"2023-07-18T03:56:15.631329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nLoading files can take a lot of time and thus I am limiting this to only the first $5354$ samples.\n","metadata":{}},{"cell_type":"code","source":"counter = 0\naudio_list = []\n\nfor x in tqdm.tqdm(os.listdir(\"/kaggle/input/bengaliai-speech/train_mp3s\") , total = 5354):\n    \n    audio , _ = librosa.load(\"/kaggle/input/bengaliai-speech/train_mp3s/\" + x)\n    audio_list.append(audio)\n    \n    counter += 1\n    if counter == 5354:break","metadata":{"execution":{"iopub.status.busy":"2023-07-18T03:57:20.248064Z","iopub.execute_input":"2023-07-18T03:57:20.248388Z","iopub.status.idle":"2023-07-18T03:57:55.890338Z","shell.execute_reply.started":"2023-07-18T03:57:20.248364Z","shell.execute_reply":"2023-07-18T03:57:55.889446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#FFA500; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #FFA500\">4 | TO DO LIST 🗂️</p>\n\n<div style=\"border-radius:10px; border:#FFA500 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\n* $TO$ $DO$ $1$ $:$ $TOKENIZE$ $SENTENCES$\n* $TO$ $DO$ $1$ $:$ $MAKE$ $DATALOADER$\n* $TO$ $DO$ $1$ $:$ $MAKE$ $A$ $MODEL$\n* $TO$ $DO$ $1$ $:$ $TRAIN$ $THE$ $MODEL$\n* $TO$ $DO$ $1$ $:$ $IMPORVE$ $REUSLTS$\n* $TO$ $DO$ $1$ $:$ $LESS$ $TRAINING$ $TIME$\n* $TO$ $DO$ $1$ $:$ $DANCE$","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#FFC0CB; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #FFC0CB\">4 | TO DO LIST 🚀</p>\n\n<div style=\"border-radius:10px; border:#FFC0CB solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\n**THAT IT FOR TODAY GUYS**\n\n**WE WILL GO DEEPER INTO THE DATA IN THE UPCOMING VERSIONS**\n\n**PLEASE COMMENT YOUR THOUGHTS, HIHGLY APPRICIATED**\n\n**DONT FORGET TO MAKE AN UPVOTE, IF YOU LIKED MY WORK $:)$**\n\n<IMG SRC = \"https://i.imgflip.com/19aadg.jpg\">\n    \n**PEACE OUT $!!!$**","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}