{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I strongly recommend to go through the article [here](https://www.analyticsvidhya.com/blog/2019/07/learn-build-first-speech-to-text-model-python/) to understand the basics of signal processing prior implementing the speech to text.\n\n**Understanding the Problem Statement for our Speech-to-Text Project**\n\nLet’s understand the problem statement of our project before we move into the implementation part.\n\nWe might be on the verge of having too many screens around us. It seems like every day, new versions of common objects are “re-invented” with built-in wifi and bright touchscreens. A promising antidote to our screen addiction is voice interfaces. \n\nTensorFlow recently released the Speech Commands Datasets. It includes 65,000 one-second long utterances of 30 short words, by thousands of different people. We’ll build a speech recognition system that understands simple spoken commands.\n\nYou can download the dataset from [here](https://www.kaggle.com/c/tensorflow-speech-recognition-challenge).\n\n**Implementing the Speech-to-Text Model in Python**\n\nThe wait is over! It’s time to build our own Speech-to-Text model from scratch.\n\n**Import the libraries**\n\nFirst, import all the necessary libraries into our notebook. LibROSA and SciPy are the Python libraries used for processing audio signals.","metadata":{}},{"cell_type":"code","source":"import os\nimport librosa\nimport IPython.display as ipd\nimport matplotlib.pyplot as plt\nimport numpy as np\nfrom scipy.io import wavfile\nimport warnings\n\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:23:15.176167Z","iopub.execute_input":"2022-06-28T09:23:15.176513Z","iopub.status.idle":"2022-06-28T09:23:17.942712Z","shell.execute_reply.started":"2022-06-28T09:23:15.176436Z","shell.execute_reply":"2022-06-28T09:23:17.941199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install py7zr","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:23:17.944499Z","iopub.execute_input":"2022-06-28T09:23:17.944785Z","iopub.status.idle":"2022-06-28T09:23:55.215917Z","shell.execute_reply.started":"2022-06-28T09:23:17.944759Z","shell.execute_reply":"2022-06-28T09:23:55.214878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from py7zr import unpack_7zarchive\nimport shutil\n\nshutil.register_unpack_format('7zip', ['.7z'], unpack_7zarchive)\nshutil.unpack_archive('/kaggle/input/tensorflow-speech-recognition-challenge/train.7z', '/kaggle/working/tensorflow-speech-recognition-challenge/')","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:23:55.217752Z","iopub.execute_input":"2022-06-28T09:23:55.218291Z","iopub.status.idle":"2022-06-28T09:32:30.430927Z","shell.execute_reply.started":"2022-06-28T09:23:55.218257Z","shell.execute_reply":"2022-06-28T09:32:30.42941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir('/kaggle/working')","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:30.433294Z","iopub.execute_input":"2022-06-28T09:32:30.433572Z","iopub.status.idle":"2022-06-28T09:32:30.44224Z","shell.execute_reply.started":"2022-06-28T09:32:30.433544Z","shell.execute_reply":"2022-06-28T09:32:30.441306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Exploration and Visualization**\n\nData Exploration and Visualization helps us to understand the data as well as pre-processing steps in a better way. \n\n**Visualization of Audio signal in time series domain**\n\nNow, we’ll visualize the audio signal in the time series domain:","metadata":{}},{"cell_type":"code","source":"train_audio_path = '/kaggle/working/tensorflow-speech-recognition-challenge/train/audio/'\nsamples, sample_rate = librosa.load(train_audio_path+'yes/0a7c2a8d_nohash_0.wav', sr = 16000)\nfig = plt.figure(figsize=(14, 8))\nax1 = fig.add_subplot(211)\nax1.set_title('Raw wave of ' + '/kaggle/working/train/audio/yes/0a7c2a8d_nohash_0.wav')\nax1.set_xlabel('time')\nax1.set_ylabel('Amplitude')\nax1.plot(np.linspace(0, sample_rate/len(samples), sample_rate), samples)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:30.443442Z","iopub.execute_input":"2022-06-28T09:32:30.444218Z","iopub.status.idle":"2022-06-28T09:32:30.679607Z","shell.execute_reply.started":"2022-06-28T09:32:30.444183Z","shell.execute_reply":"2022-06-28T09:32:30.678341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Sampling rate **\n\nLet us now look at the sampling rate of the audio signals","metadata":{}},{"cell_type":"code","source":"ipd.Audio(samples, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:30.680808Z","iopub.execute_input":"2022-06-28T09:32:30.68114Z","iopub.status.idle":"2022-06-28T09:32:30.69705Z","shell.execute_reply.started":"2022-06-28T09:32:30.681106Z","shell.execute_reply":"2022-06-28T09:32:30.696328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(sample_rate)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:30.698439Z","iopub.execute_input":"2022-06-28T09:32:30.699275Z","iopub.status.idle":"2022-06-28T09:32:30.703449Z","shell.execute_reply.started":"2022-06-28T09:32:30.699243Z","shell.execute_reply":"2022-06-28T09:32:30.702719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Resampling**\n\nFrom the above, we can understand that the sampling rate of the signal is 16000 hz. Let us resample it to 8000 hz since most of the speech related frequencies are present in 8000z ","metadata":{}},{"cell_type":"code","source":"samples = librosa.resample(samples, sample_rate, 8000)\nipd.Audio(samples, rate=8000)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:30.704579Z","iopub.execute_input":"2022-06-28T09:32:30.705568Z","iopub.status.idle":"2022-06-28T09:32:31.531948Z","shell.execute_reply.started":"2022-06-28T09:32:30.705516Z","shell.execute_reply":"2022-06-28T09:32:31.531103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let’s understand the number of recordings for each voice command:","metadata":{}},{"cell_type":"code","source":"labels=os.listdir(train_audio_path)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:31.533155Z","iopub.execute_input":"2022-06-28T09:32:31.533481Z","iopub.status.idle":"2022-06-28T09:32:31.537813Z","shell.execute_reply.started":"2022-06-28T09:32:31.533453Z","shell.execute_reply":"2022-06-28T09:32:31.536909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#find count of each label and plot bar graph\nno_of_recordings=[]\nfor label in labels:\n    waves = [f for f in os.listdir(train_audio_path + '/'+ label) if f.endswith('.wav')]\n    no_of_recordings.append(len(waves))\n    \n#plot\nplt.figure(figsize=(30,5))\nindex = np.arange(len(labels))\nplt.bar(index, no_of_recordings)\nplt.xlabel('Commands', fontsize=12)\nplt.ylabel('No of recordings', fontsize=12)\nplt.xticks(index, labels, fontsize=15, rotation=60)\nplt.title('No. of recordings for each command')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:31.542404Z","iopub.execute_input":"2022-06-28T09:32:31.543037Z","iopub.status.idle":"2022-06-28T09:32:31.942262Z","shell.execute_reply.started":"2022-06-28T09:32:31.54299Z","shell.execute_reply":"2022-06-28T09:32:31.941379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels=[\"yes\", \"no\", \"up\", \"down\", \"left\", \"right\", \"on\", \"off\", \"stop\", \"go\"]","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:31.943134Z","iopub.execute_input":"2022-06-28T09:32:31.943766Z","iopub.status.idle":"2022-06-28T09:32:31.947275Z","shell.execute_reply.started":"2022-06-28T09:32:31.943741Z","shell.execute_reply":"2022-06-28T09:32:31.946553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(labels))","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:31.948176Z","iopub.execute_input":"2022-06-28T09:32:31.948441Z","iopub.status.idle":"2022-06-28T09:32:31.961998Z","shell.execute_reply.started":"2022-06-28T09:32:31.948417Z","shell.execute_reply":"2022-06-28T09:32:31.961189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Duration of recordings**\n\nWhat’s next? A look at the distribution of the duration of recordings:","metadata":{}},{"cell_type":"code","source":"duration_of_recordings=[]\nfor label in labels:\n    waves = [f for f in os.listdir(train_audio_path + '/'+ label) if f.endswith('.wav')]\n    for wav in waves:\n        sample_rate, samples = wavfile.read(train_audio_path + '/' + label + '/' + wav)\n        duration_of_recordings.append(float(len(samples)/sample_rate))\n    \nplt.hist(np.array(duration_of_recordings))","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:31.963039Z","iopub.execute_input":"2022-06-28T09:32:31.963885Z","iopub.status.idle":"2022-06-28T09:32:33.352847Z","shell.execute_reply.started":"2022-06-28T09:32:31.963852Z","shell.execute_reply":"2022-06-28T09:32:33.352211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Preprocessing the audio waves**\n\nIn the data exploration part earlier, we have seen that the duration of a few recordings is less than 1 second and the sampling rate is too high. So, let us read the audio waves and use the below-preprocessing steps to deal with this.\n\nHere are the two steps we’ll follow:\n\n* Resampling\n* Removing shorter commands of less than 1 second\n\nLet us define these preprocessing steps in the below code snippet:","metadata":{}},{"cell_type":"code","source":"train_audio_path = '/kaggle/working/tensorflow-speech-recognition-challenge/train/audio/'\n\nall_wave = []\nall_label = []\nfor label in labels:\n    print(label)\n    waves = [f for f in os.listdir(train_audio_path + '/'+ label) if f.endswith('.wav')]\n    for wav in waves:\n        samples, sample_rate = librosa.load(train_audio_path + '/' + label + '/' + wav, sr = 16000)\n        samples = librosa.resample(samples, sample_rate, 8000)\n        if(len(samples)== 8000) : \n            all_wave.append(samples)\n            all_label.append(label)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:32:33.35405Z","iopub.execute_input":"2022-06-28T09:32:33.354476Z","iopub.status.idle":"2022-06-28T09:38:14.724912Z","shell.execute_reply.started":"2022-06-28T09:32:33.354451Z","shell.execute_reply":"2022-06-28T09:38:14.723927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Convert the output labels to integer encoded:","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\nle = LabelEncoder()\ny=le.fit_transform(all_label)\nclasses= list(le.classes_)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:14.72581Z","iopub.execute_input":"2022-06-28T09:38:14.726118Z","iopub.status.idle":"2022-06-28T09:38:14.738161Z","shell.execute_reply.started":"2022-06-28T09:38:14.726063Z","shell.execute_reply":"2022-06-28T09:38:14.737342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, convert the integer encoded labels to a one-hot vector since it is a multi-classification problem:","metadata":{}},{"cell_type":"code","source":"from keras.utils import np_utils\ny=np_utils.to_categorical(y, num_classes=len(labels))","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:14.739126Z","iopub.execute_input":"2022-06-28T09:38:14.739403Z","iopub.status.idle":"2022-06-28T09:38:21.429491Z","shell.execute_reply.started":"2022-06-28T09:38:14.739369Z","shell.execute_reply":"2022-06-28T09:38:21.42849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reshape the 2D array to 3D since the input to the conv1d must be a 3D array:","metadata":{}},{"cell_type":"code","source":"all_wave = np.array(all_wave).reshape(-1,8000,1)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:21.430896Z","iopub.execute_input":"2022-06-28T09:38:21.4319Z","iopub.status.idle":"2022-06-28T09:38:21.641615Z","shell.execute_reply.started":"2022-06-28T09:38:21.431859Z","shell.execute_reply":"2022-06-28T09:38:21.640609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Split into train and validation set**\n\nNext, we will train the model on 80% of the data and validate on the remaining 20%:\n","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nx_tr, x_val, y_tr, y_val = train_test_split(np.array(all_wave),np.array(y),stratify=y,test_size = 0.2,random_state=777,shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:21.645641Z","iopub.execute_input":"2022-06-28T09:38:21.645936Z","iopub.status.idle":"2022-06-28T09:38:22.230901Z","shell.execute_reply.started":"2022-06-28T09:38:21.645907Z","shell.execute_reply":"2022-06-28T09:38:22.230146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Model Architecture for this problem**\n\nWe will build the speech-to-text model using conv1d. Conv1d is a convolutional neural network which performs the convolution along only one dimension. ","metadata":{}},{"cell_type":"markdown","source":"**Model building**\n\nLet us implement the model using Keras functional API.","metadata":{}},{"cell_type":"code","source":"from keras.layers import Dense, Dropout, Flatten, Conv1D, Input, MaxPooling1D\nfrom keras.models import Model\nfrom keras.callbacks import EarlyStopping, ModelCheckpoint\nfrom keras import backend as K\nK.clear_session()\n\ninputs = Input(shape=(8000,1))\n\n#First Conv1D layer\nconv = Conv1D(8,13, padding='valid', activation='relu', strides=1)(inputs)\nconv = MaxPooling1D(3)(conv)\nconv = Dropout(0.3)(conv)\n\n#Second Conv1D layer\nconv = Conv1D(16, 11, padding='valid', activation='relu', strides=1)(conv)\nconv = MaxPooling1D(3)(conv)\nconv = Dropout(0.3)(conv)\n\n#Third Conv1D layer\nconv = Conv1D(32, 9, padding='valid', activation='relu', strides=1)(conv)\nconv = MaxPooling1D(3)(conv)\nconv = Dropout(0.3)(conv)\n\n#Fourth Conv1D layer\nconv = Conv1D(64, 7, padding='valid', activation='relu', strides=1)(conv)\nconv = MaxPooling1D(3)(conv)\nconv = Dropout(0.3)(conv)\n\n#Flatten layer\nconv = Flatten()(conv)\n\n#Dense Layer 1\nconv = Dense(256, activation='relu')(conv)\nconv = Dropout(0.3)(conv)\n\n#Dense Layer 2\nconv = Dense(128, activation='relu')(conv)\nconv = Dropout(0.3)(conv)\n\noutputs = Dense(len(labels), activation='softmax')(conv)\n\nmodel = Model(inputs, outputs)\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:22.2323Z","iopub.execute_input":"2022-06-28T09:38:22.232652Z","iopub.status.idle":"2022-06-28T09:38:22.4593Z","shell.execute_reply.started":"2022-06-28T09:38:22.232614Z","shell.execute_reply":"2022-06-28T09:38:22.458431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Define the loss function to be categorical cross-entropy since it is a multi-classification problem:","metadata":{}},{"cell_type":"code","source":"model.compile(loss='categorical_crossentropy',optimizer='adam',metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:22.460471Z","iopub.execute_input":"2022-06-28T09:38:22.460743Z","iopub.status.idle":"2022-06-28T09:38:22.474001Z","shell.execute_reply.started":"2022-06-28T09:38:22.460717Z","shell.execute_reply":"2022-06-28T09:38:22.473342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Early stopping and model checkpoints are the callbacks to stop training the neural network at the right time and to save the best model after every epoch:","metadata":{}},{"cell_type":"code","source":"metric='val_accuracy'\n\nes = EarlyStopping(monitor='val_loss', mode='min', verbose=1, patience=10, min_delta=0.001) \nmc = ModelCheckpoint('/kaggle/working/best_model.hdf5', monitor=metric, verbose=1, save_best_only=True, mode='max')","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:22.475369Z","iopub.execute_input":"2022-06-28T09:38:22.475746Z","iopub.status.idle":"2022-06-28T09:38:22.485959Z","shell.execute_reply.started":"2022-06-28T09:38:22.475708Z","shell.execute_reply":"2022-06-28T09:38:22.485226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us train the model on a batch size of 32 and evaluate the performance on the holdout set:","metadata":{}},{"cell_type":"code","source":"history=model.fit(x_tr, y_tr ,epochs=100, callbacks=[es,mc], batch_size=32, validation_data=(x_val,y_val))","metadata":{"execution":{"iopub.status.busy":"2022-06-28T09:38:22.488899Z","iopub.execute_input":"2022-06-28T09:38:22.490152Z","iopub.status.idle":"2022-06-28T10:04:46.099254Z","shell.execute_reply.started":"2022-06-28T09:38:22.4901Z","shell.execute_reply":"2022-06-28T10:04:46.098162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pwd","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:04:46.100729Z","iopub.execute_input":"2022-06-28T10:04:46.101743Z","iopub.status.idle":"2022-06-28T10:04:46.109034Z","shell.execute_reply.started":"2022-06-28T10:04:46.101706Z","shell.execute_reply":"2022-06-28T10:04:46.107868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ls","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:04:46.110398Z","iopub.execute_input":"2022-06-28T10:04:46.11118Z","iopub.status.idle":"2022-06-28T10:04:46.487488Z","shell.execute_reply.started":"2022-06-28T10:04:46.111145Z","shell.execute_reply":"2022-06-28T10:04:46.486412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Diagnostic plot**\n\nI’m going to lean on visualization again to understand the performance of the model over a period of time:","metadata":{}},{"cell_type":"code","source":"from matplotlib import pyplot\npyplot.plot(history.history['loss'], label='train')\npyplot.plot(history.history['val_loss'], label='test')\npyplot.legend()\npyplot.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:04:46.488516Z","iopub.execute_input":"2022-06-28T10:04:46.488752Z","iopub.status.idle":"2022-06-28T10:05:46.465811Z","shell.execute_reply.started":"2022-06-28T10:04:46.488728Z","shell.execute_reply":"2022-06-28T10:05:46.465255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Loading the best model**","metadata":{}},{"cell_type":"code","source":"from keras.models import load_model\nmodel=load_model('best_model.hdf5')","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:05:46.466765Z","iopub.execute_input":"2022-06-28T10:05:46.467291Z","iopub.status.idle":"2022-06-28T10:05:46.661958Z","shell.execute_reply.started":"2022-06-28T10:05:46.467262Z","shell.execute_reply":"2022-06-28T10:05:46.661076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Define the function that predicts text for the given audio:","metadata":{}},{"cell_type":"code","source":"def predict(audio):\n    prob=model.predict(audio.reshape(1,8000,1))\n    index=np.argmax(prob[0])\n    return classes[index]","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:05:46.662949Z","iopub.execute_input":"2022-06-28T10:05:46.663201Z","iopub.status.idle":"2022-06-28T10:05:46.668586Z","shell.execute_reply.started":"2022-06-28T10:05:46.663177Z","shell.execute_reply":"2022-06-28T10:05:46.667446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Prediction time! Make predictions on the validation data:","metadata":{}},{"cell_type":"code","source":"import random\nindex=random.randint(0,len(x_val)-1)\nsamples=x_val[index].ravel()\nprint(\"Audio:\",classes[np.argmax(y_val[index])])\nipd.Audio(samples, rate=8000)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:05:46.672725Z","iopub.execute_input":"2022-06-28T10:05:46.673036Z","iopub.status.idle":"2022-06-28T10:05:46.686376Z","shell.execute_reply.started":"2022-06-28T10:05:46.673008Z","shell.execute_reply":"2022-06-28T10:05:46.685412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Text:\",predict(samples))","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:05:46.68761Z","iopub.execute_input":"2022-06-28T10:05:46.687888Z","iopub.status.idle":"2022-06-28T10:05:46.896959Z","shell.execute_reply.started":"2022-06-28T10:05:46.687865Z","shell.execute_reply":"2022-06-28T10:05:46.896143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The best part is yet to come! Here is a script that prompts a user to record voice commands. Record your own voice commands and test it on the model:","metadata":{}},{"cell_type":"code","source":"!pip install sounddevice\n!sudo apt-get install libportaudio2","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:05:46.898195Z","iopub.execute_input":"2022-06-28T10:05:46.898474Z","iopub.status.idle":"2022-06-28T10:06:02.099294Z","shell.execute_reply.started":"2022-06-28T10:05:46.898447Z","shell.execute_reply":"2022-06-28T10:06:02.098546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sounddevice as sd\nimport soundfile as sf\n\nsamplerate = 16000\nduration = 1 # seconds\nfilename = 'yes.wav'\nprint(\"start\")\nmydata = sd.rec(int(samplerate * duration), samplerate=samplerate,\n    channels=1, blocking=True)\nprint(\"end\")\nsd.wait()\nsf.write(filename, mydata, samplerate)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:06:02.100372Z","iopub.execute_input":"2022-06-28T10:06:02.100732Z","iopub.status.idle":"2022-06-28T10:06:02.570671Z","shell.execute_reply.started":"2022-06-28T10:06:02.100698Z","shell.execute_reply":"2022-06-28T10:06:02.569415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us now read the saved voice command and convert it to text:","metadata":{}},{"cell_type":"code","source":"os.listdir('/kaggle/working/voice-commands/prateek_voice_v2')","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:06:02.571566Z","iopub.status.idle":"2022-06-28T10:06:02.571856Z","shell.execute_reply.started":"2022-06-28T10:06:02.571723Z","shell.execute_reply":"2022-06-28T10:06:02.571739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filepath='/kaggle/working/voice-commands/prateek_voice_v2'","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:06:02.572524Z","iopub.status.idle":"2022-06-28T10:06:02.572796Z","shell.execute_reply.started":"2022-06-28T10:06:02.572665Z","shell.execute_reply":"2022-06-28T10:06:02.572679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#reading the voice commands\nsamples, sample_rate = librosa.load(filepath + '/' + 'stop.wav', sr = 16000)\nsamples = librosa.resample(samples, sample_rate, 8000)\nipd.Audio(samples,rate=8000)              ","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:06:02.573877Z","iopub.status.idle":"2022-06-28T10:06:02.574184Z","shell.execute_reply.started":"2022-06-28T10:06:02.574022Z","shell.execute_reply":"2022-06-28T10:06:02.574036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#converting voice commands to text\npredict(samples)","metadata":{"execution":{"iopub.status.busy":"2022-06-28T10:06:02.575534Z","iopub.status.idle":"2022-06-28T10:06:02.576225Z","shell.execute_reply.started":"2022-06-28T10:06:02.576048Z","shell.execute_reply":"2022-06-28T10:06:02.576064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Congratulations! You have just built your very own speech-to-text model!","metadata":{}}]}