{"metadata":{"kernelspec":{"language":"python","name":"python3","display_name":"Python 3"},"language_info":{"name":"python","file_extension":".py","codemirror_mode":{"version":3,"name":"ipython"},"version":"3.6.3","nbconvert_exporter":"python","mimetype":"text/x-python","pygments_lexer":"ipython3"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Introduction \n\nHello! This is my very first Kernel. It is meant to give a grasp of a problem of speech representation. I'd also like to take a look on a features specific to this dataset. \n\nContent:<br>\n* [1. Visualization of the recordings - input features](#visualization)\n   * [1.1. Wave and spectrogram](#waveandspectrogram)\n   * [1.2. MFCC](#mfcc)\n   * [1.3. Sprectrogram in 3d](#3d)\n   * [1.4. Silence removal](#resampl)\n   * [1.5. Resampling - dimensionality reductions](#silenceremoval)\n   * [1.6. Features extraction steps](#featuresextractionsteps)\n* [2. Dataset investigation](#investigations)\n   * [2.1. Number of files](#numberoffiles)\n   * [2.2. Mean spectrograms and fft](#meanspectrogramsandfft)\n   * [2.3. Deeper into recordings](#deeper)\n   * [2.4. Length of recordings](#len)\n   * [2.5. Note on Gaussian Mixtures modeling](#gmms)\n   * [2.6. Frequency components across the words](#components)\n   * [2.7. Anomaly detection](#anomaly)\n* [3. Where to look for the inspiration](#wheretostart)\n\nAll we need is here:","metadata":{"_uuid":"cf808b6b464476e44c8dda4beada9884cf91adca","_cell_guid":"cec0e376-8e36-4751-8b53-2f298fda3a47"}},{"cell_type":"markdown","source":"データセットの説明\n以下は、このコンテストのデータの概要です。 「概要」ページのタブにある補足資料を読むことを強くお勧めします。\n\nファイルの説明\ntrain.7z - いくつかの情報ファイルとオーディオ ファイルのフォルダーが含まれています。 オーディオ フォルダーには、音声コマンドの 1 秒クリップを含むサブフォルダーが含まれており、フォルダー名はオーディオ クリップのラベルです。 予測すべきラベルは他にもあります。 テストで予測する必要があるラベルは、yes、no、up、down、left、right、on、off、stop、go です。 それ以外のものはすべて未知か沈黙のいずれかであると考えるべきです。 _background_noise_ フォルダーには、分割してトレーニング入力として使用できる「沈黙」の長いクリップが含まれています。\n\n★zipファイルの解凍がうまくいかなかったので、別のデータセットを取り込んで実行した。\nhttps://www.kaggle.com/datasets/guntherneumair1/speechcommandv02-cleaned\nただし、good_files, bad_filesの区分になっていたり、一部のファイルが削除？されている模様。\n\n\nコンテクスト：元の音声コマンド データセットはクラウドソースで作成されたもので、カットオフ、無音、および間違ったサンプルが多数含まれています。これにより、最大パフォーマンスが確実に制限されます。\nコンテンツ：トレーニングされたモデルを使用して、誤認識されたすべてのファイルをフォルダーにプッシュし、手動で確認しました。\n\n\nトレーニング オーディオに含まれるファイルは、ラベル間で一意の名前が付けられませんが、ラベル フォルダーを含めると一意になります。 たとえば、00f0204f_nohash_0.wav は 14 個のフォルダーにありますが、そのファイルはフォルダーごとに異なる音声コマンドです。\n\nファイルには、最初の要素が音声コマンドを与えた人の件名 ID になり、最後の要素が繰り返しのコマンドを示すように名前が付けられます。 反復コマンドとは、被験者が同じ単語を複数回繰り返すことです。 テスト データには被験者 ID が提供されていないため、テスト データ内のコマンドの大部分は電車内では見られない被験者からのものであると推測できます。\n\nトレーニング データのプロパティ (音声の長さなど) に不一致が生じることが予想されます。\n\ntest.7z - Clip_000044442.wav 形式の 150,000 以上のファイルを含むオーディオ フォルダーが含まれています。 タスクは、正しいラベルを予測することです。 すべてのファイルがリーダーボード スコアとして評価されるわけではありません。\n\nSample_submission.csv - 正しい形式のサンプル送信ファイル。\n\nlink_to_gcp_credits_form.txt - 最初の 500 人のリクエスト者に提供される、500 ドルの GCP クレジットをリクエストするための URL を提供します。 単位認定に関するこちらの説明をご覧ください。\n\n以下は、Train フォルダーに含まれる README であり、より詳細な情報が含まれています。\n\n音声コマンド データ セット v0.01\nこれは 1 秒の .wav オーディオ ファイルのセットで、それぞれに 1 つの音声ファイルが含まれています。\n英単語。 これらの言葉は小さなコマンドセットからのものであり、\nさまざまなスピーカー。 オーディオ ファイルは、次のようなフォルダーに分類されます。\nこのデータセットは、単純な単語のトレーニングに役立つように設計されています。\n機械学習モデル。\n\nCreative Commons BY 4.0 ライセンスに基づいてライセンスされています。 ライセンスを参照\n詳細については、このフォルダー内のファイルを参照してください。 元の場所は次のとおりです\nhttp://download.tensorflow.org/data/speech_commands_v0.01.tar.gz。\n\n歴史\nこれは、64,727 個のオーディオ ファイルを含むデータ セットのバージョン 0.01 で、\n2017年8月3日。\n\nコレクション\n音声ファイルはクラウドソーシングを使用して収集されました。を参照してください。\naiyprojects.withgoogle.com/open_speech_recording\n私たちが使用したオープンソースのオーディオ コレクション コードの一部については、次のことを考慮してください。\nこのデータセットの拡大に貢献しています)。 目的は、例を収集することでした。\n人々は会話文ではなく、単一単語のコマンドを話すので、\n彼らは 5 分間にわたって個々の単語を入力するよう求められました\nセッション。 20 の主要なコマンドワードが記録され、ほとんどの話者がそれぞれの言葉を発していました。\nそのうち5回。 中心となる単語は\"Yes\", \"No\", \"Up\", \"Down\", \"Left\",\n\"Right\", \"On\", \"Off\", \"Stop\", \"Go\", \"Zero\", \"One\", \"Two\", \"Three\", \"Four\",\n\"Five\", \"Six\", \"Seven\", \"Eight\", \"Nine\"。 認識されないものを区別しやすくするため\n単語に加えて、ほとんどの話者が 1 回しか言わない補助単語も 10 個あります。\nこれらには、\"Bed\", \"Bird\", \"Cat\", \"Dog\", \"Happy\", \"House\", \"Marvin\", \"Sheila\",\n\"Tree\", and \"Wow\"。\n\n組織\nファイルはフォルダーに編成され、各ディレクトリ名にラベルが付けられます。\n含まれているすべての音声ファイルで話されている単語。 詳細は残されていなかった\n参加者の年齢、性別、居住地、およびランダムな ID が割り当てられました。\n一人ひとりに。 ただし、これらの ID は安定しており、各ファイル名にエンコードされています。\nアンダースコアの前の最初の部分として。 参加者が複数回寄付した場合\n同じ単語の発話は、末尾の番号によって区別されます。\nファイル名。 たとえば、ファイル パス happy/3cfc6b3a_nohash_2.wav\n話された単語が「ハッピー」であり、話者の ID が「3cfc6b3a」であることを示します。\nこれは、データセット内でこの話者によるその単語の 3 番目の発話です。 の\n「nohash」セクションは、1 人の話者によるすべての発話が確実に行われるようにするためのものです。\n非常に類似した繰り返しを避けるために、同じトレーニング パーティションに分類されます。\n非現実的に楽観的な評価スコアを与える。\n\nパーティショニング\nオーディオ クリップはトレーニング、テスト、検証セットに分割されていません。\n明示的に、ただし慣例により、それぞれを安定して割り当てるためにハッシュ関数が使用されます。\nファイルをセットに追加します。 以下は、完全なファイル パスを取得する方法を示す Python コードです。\nそして、必要な検証セットとテストセットのサイズ (通常は両方とも 10%) を使用して、\nセットを割り当てる：\n\nDataset Description  \nThe following is a high level overview of the data for this competition. It is highly recommended that you read the supplementary materials found on the tabs of the Overview page.\n\nFile descriptions  \ntrain.7z - Contains a few informational files and a folder of audio files. The audio folder contains subfolders with 1 second clips of voice commands, with the folder name being the label of the audio clip. There are more labels that should be predicted. The labels you will need to predict in Test are yes, no, up, down, left, right, on, off, stop, go. Everything else should be considered either unknown or silence. The folder _background_noise_ contains longer clips of \"silence\" that you can break up and use as training input.\n\nThe files contained in the training audio are not uniquely named across labels, but they are unique if you include the label folder. For example, 00f0204f_nohash_0.wav is found in 14 folders, but that file is a different speech command in each folder.\n\nThe files are named so the first element is the subject id of the person who gave the voice command, and the last element indicated repeated commands. Repeated commands are when the subject repeats the same word multiple times. Subject id is not provided for the test data, and you can assume that the majority of commands in the test data were from subjects not seen in train.\n\nYou can expect some inconsistencies in the properties of the training data (e.g., length of the audio).\n\ntest.7z - Contains an audio folder with 150,000+ files in the format clip_000044442.wav. The task is to predict the correct label. Not all of the files are evaluated for the leaderboard score.\n\nsample_submission.csv - A sample submission file in the correct format.\n\nlink_to_gcp_credits_form.txt - Provides the URL to request $500 in GCP credits, provided to the first 500 requestors. Please see this clarification on credit qualification.\n\nThe following is the README contained in the Train folder, and contains more detailed information.\n\nSpeech Commands Data Set v0.01  \nThis is a set of one-second .wav audio files, each containing a single spoken\nEnglish word. These words are from a small set of commands, and are spoken by a\nvariety of different speakers. The audio files are organized into folders based\non the word they contain, and this data set is designed to help train simple\nmachine learning models.\n\nIt's licensed under the Creative Commons BY 4.0 license. See the LICENSE\nfile in this folder for full details. Its original location was at\nhttp://download.tensorflow.org/data/speech_commands_v0.01.tar.gz.\n\nHistory  \nThis is version 0.01 of the data set containing 64,727 audio files, released on\nAugust 3rd 2017.\n\nCollection  \nThe audio files were collected using crowdsourcing, see\naiyprojects.withgoogle.com/open_speech_recording\nfor some of the open source audio collection code we used (and please consider\ncontributing to enlarge this data set). The goal was to gather examples of\npeople speaking single-word commands, rather than conversational sentences, so\nthey were prompted for individual words over the course of a five minute\nsession. Twenty core command words were recorded, with most speakers saying each\nof them five times. The core words are \"Yes\", \"No\", \"Up\", \"Down\", \"Left\",\n\"Right\", \"On\", \"Off\", \"Stop\", \"Go\", \"Zero\", \"One\", \"Two\", \"Three\", \"Four\",\n\"Five\", \"Six\", \"Seven\", \"Eight\", and \"Nine\". To help distinguish unrecognized\nwords, there are also ten auxiliary words, which most speakers only said once.\nThese include \"Bed\", \"Bird\", \"Cat\", \"Dog\", \"Happy\", \"House\", \"Marvin\", \"Sheila\",\n\"Tree\", and \"Wow\".\n\nOrganization  \nThe files are organized into folders, with each directory name labelling the\nword that is spoken in all the contained audio files. No details were kept of\nany of the participants age, gender, or location, and random ids were assigned\nto each individual. These ids are stable though, and encoded in each file name\nas the first part before the underscore. If a participant contributed multiple\nutterances of the same word, these are distinguished by the number at the end of\nthe file name. For example, the file path happy/3cfc6b3a_nohash_2.wav\nindicates that the word spoken was \"happy\", the speaker's id was \"3cfc6b3a\", and\nthis is the third utterance of that word by this speaker in the data set. The\n'nohash' section is to ensure that all the utterances by a single speaker are\nsorted into the same training partition, to keep very similar repetitions from\ngiving unrealistically optimistic evaluation scores.\n\nPartitioning  \nThe audio clips haven't been separated into training, test, and validation sets\nexplicitly, but by convention a hashing function is used to stably assign each\nfile to a set. Here's some Python code demonstrating how a complete file path\nand the desired validation and test set sizes (usually both 10%) are used to\nassign a set:","metadata":{}},{"cell_type":"code","source":"MAX_NUM_WAVS_PER_CLASS = 2**27 - 1  # ~134M\n\ndef which_set(filename, validation_percentage, testing_percentage):\n  \"\"\"ファイルがどのデータ パーティションに属するかを決定します。\n\n   ファイルを同じトレーニング、検証、またはテスト セット内に保持したいと考えています。\n   時間の経過とともに新しいものが追加される場合。 これにより、テストが行われる可能性が低くなります\n   長時間の実行が再開されると、サンプルがトレーニングで誤って再利用されてしまう\n   例えば。 この安定性を維持するために、ファイル名のハッシュが取得されて使用されます。\n   どのセットに属するかを決定します。 この決定は以下にのみ依存します。\n   名前と設定された比率なので、他のファイルが追加されても変更されません。\n\n   特定のファイルを関連ファイルとして関連付けることも便利です (単語など)。\n   同じ人が話したもの)、ファイル名の「_nohash_」以降はすべて\n   セットの決定では無視されます。 これにより、「bobby_nohash_0.wav」と\n   たとえば、「bobby_nohash_1.wav」は常に同じセット内にあります。\n\n   引数:\n     filename: データ サンプルのファイル パス。\n     validation_percentage: 検証に使用するデータセットの量。\n     testing_percentage: テストに使用するデータセットの量。\n\n   戻り値：\n     「トレーニング」、「検証」、または「テスト」のいずれかの文字列。\n     \n     Determines which data partition the file should belong to.\n\n  We want to keep files in the same training, validation, or testing sets even\n  if new ones are added over time. This makes it less likely that testing\n  samples will accidentally be reused in training when long runs are restarted\n  for example. To keep this stability, a hash of the filename is taken and used\n  to determine which set it should belong to. This determination only depends on\n  the name and the set proportions, so it won't change as other files are added.\n\n  It's also useful to associate particular files as related (for example words\n  spoken by the same person), so anything after '_nohash_' in a filename is\n  ignored for set determination. This ensures that 'bobby_nohash_0.wav' and\n  'bobby_nohash_1.wav' are always in the same set, for example.\n\n  Args:\n    filename: File path of the data sample.\n    validation_percentage: How much of the data set to use for validation.\n    testing_percentage: How much of the data set to use for testing.\n\n  Returns:\n    String, one of 'training', 'validation', or 'testing'.\n  \"\"\"\n  base_name = os.path.basename(filename)\n  # We want to ignore anything after '_nohash_' in the file name when\n  # deciding which set to put a wav in, so the data set creator has a way of\n  # grouping wavs that are close variations of each other.\n  hash_name = re.sub(r'_nohash_.*$', '', base_name)\n  # This looks a bit magical, but we need to decide whether this file should\n  # go into the training, testing, or validation sets, and we want to keep\n  # existing files in the same set even if more files are subsequently\n  # added.\n  # To do that, we need a stable way of deciding based on just the file name\n  # itself, so we do a hash of that and then use that to generate a\n  # probability value that we use to assign it.\n  hash_name_hashed = hashlib.sha1(hash_name).hexdigest()\n  percentage_hash = ((int(hash_name_hashed, 16) %\n                      (MAX_NUM_WAVS_PER_CLASS + 1)) *\n                     (100.0 / MAX_NUM_WAVS_PER_CLASS))\n  if percentage_hash < validation_percentage:\n    result = 'validation'\n  elif percentage_hash < (testing_percentage + validation_percentage):\n    result = 'testing'\n  else:\n    result = 'training'\n  return result","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:36.705391Z","iopub.execute_input":"2023-09-09T23:24:36.706163Z","iopub.status.idle":"2023-09-09T23:24:36.733077Z","shell.execute_reply.started":"2023-09-09T23:24:36.706075Z","shell.execute_reply":"2023-09-09T23:24:36.732374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"現在のセットに対してこれを実行した結果は、このアーカイブに次のように含まれます。\nvalidation_list.txt と testing_list.txt。 これらのテキスト ファイルには、次のパスが含まれています。\n各セット内のすべてのファイルを、新しい行に各パスを入れて表示します。 そうでないファイル\nこれらのリストのいずれかにあるものは、トレーニング セットの一部と見なすことができます。\n\n処理  \nオリジナルの音声ファイルは人々によって管理されていない場所で収集されました\n世界中で。 収録は密室でお願いしました。\nプライバシー上の理由はありましたが、品質要件は規定されていませんでした。 これは、\n私たちが考えているような音声データの例が欲しかったので、設計しました。\n消費者向けアプリケーションやロボティクスアプリケーションで遭遇することは、私たちがあまり経験していないことです。\n録音機器や環境の制御。 データは次の形式でキャプチャされました。\nさまざまな形式 (Web アプリの Ogg Vorbis エンコーディングなど)\n16000 サンプルで 16 ビット リトル エンディアン PCM エンコードされた WAVE ファイルに変換\nレート。 次に、音声はほとんどの部分を揃えるために 1 秒の長さにトリミングされました。\nを使用した発話\nextract_loudest_section\n道具。 その後、音声ファイルは沈黙や間違った単語がないかスクリーニングされ、\nラベルごとにフォルダーに整理されます。\n\n背景雑音  \nノイズの多い環境に対処できるようにネットワークを訓練するには、次のことを組み合わせると効果的です。\nリアルなバックグラウンドオーディオで。 _background_noise_ フォルダーには、次のセットが含まれています。\n録音または数学的シミュレーションである長いオーディオ クリップ\nノイズ。 詳細については、_background_noise_/README.md を参照してください。\n\n引用  \n仕事で Speech Commands データセットを使用する場合は、次のように引用してください。\n\nAPA スタイルの引用: 「Warden P. Speech Commands: 単一単語の公開データセット」\n音声認識、2017 年。\nhttp://download.tensorflow.org/data/speech_commands_v0.01.tar.gz」。\n\nBibTeX @article{speechcommands, title={音声コマンド: 単一単語音声認識用の公開データセット。}、author={Warden、Pete}、journal={http://download.tensorflow.org/data/ から入手可能なデータセット speech_commands_v0.01.tar.gz}、年={2017} }\n\nクレジット  \nこのデータセットに録音を提供してくださった皆様に多大な感謝を申し上げます。\nとてもありがたい。 私も助けがなければこれを組み立てることはできませんでした\nビリー・ラトリッジ、ラジャット・モンガ、ラジエル・アルバレス、ブラッド・クルーガー、バーバラのサポート\nPetit、Gursheesh Kour、AIY チームと TensorFlow チームの皆様。\n\nピート ウォーデン、petewarden@google.com\n\nThe results of running this over the current set are included in this archive as\nvalidation_list.txt and testing_list.txt. These text files contain the paths to\nall the files in each set, with each path on a new line. Any files that aren't\nin either of these lists can be considered to be part of the training set.\n\nProcessing  \nThe original audio files were collected in uncontrolled locations by people\naround the world. We requested that they do the recording in a closed room for\nprivacy reasons, but didn't stipulate any quality requirements. This was by\ndesign, since we wanted examples of the sort of speech data that we're likely to\nencounter in consumer and robotics applications, where we don't have much\ncontrol over the recording equipment or environment. The data was captured in a\nvariety of formats, for example Ogg Vorbis encoding for the web app, and then\nconverted to a 16-bit little-endian PCM-encoded WAVE file at a 16000 sample\nrate. The audio was then trimmed to a one second length to align most\nutterances, using the\nextract_loudest_section\ntool. The audio files were then screened for silence or incorrect words, and\narranged into folders by label.\n\nBackground Noise  \nTo help train networks to cope with noisy environments, it can be helpful to mix\nin realistic background audio. The _background_noise_ folder contains a set of\nlonger audio clips that are either recordings or mathematical simulations of\nnoise. For more details, see the _background_noise_/README.md.\n\nCitations  \nIf you use the Speech Commands dataset in your work, please cite it as:\n\nAPA-style citation: \"Warden P. Speech Commands: A public dataset for single-word\nspeech recognition, 2017. Available from\nhttp://download.tensorflow.org/data/speech_commands_v0.01.tar.gz\".\n\nBibTeX @article{speechcommands, title={Speech Commands: A public dataset for single-word speech recognition.}, author={Warden, Pete}, journal={Dataset available from http://download.tensorflow.org/data/speech_commands_v0.01.tar.gz}, year={2017} }\n\nCredits  \nMassive thanks are due to everyone who donated recordings to this data set, I'm\nvery grateful. I also couldn't have put this together without the help and\nsupport of Billy Rutledge, Rajat Monga, Raziel Alvarez, Brad Krueger, Barbara\nPetit, Gursheesh Kour, and all the AIY and TensorFlow teams.\n\nPete Warden, petewarden@google.com  ","metadata":{}},{"cell_type":"code","source":"import os\nfrom os.path import isdir, join\nfrom pathlib import Path\nimport pandas as pd\n\n# Math\nimport numpy as np\nfrom scipy.fftpack import fft\nfrom scipy import signal\nfrom scipy.io import wavfile\nimport librosa\n\nfrom sklearn.decomposition import PCA\n\n# Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport IPython.display as ipd\nimport librosa.display\n\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.tools as tls\nimport pandas as pd\n\n%matplotlib inline","metadata":{"_uuid":"d9596d80cb6445d4214dda15e40d777cadbd4669","_cell_guid":"8fd82027-7be0-4a4e-a921-b8acacaaf077","execution":{"iopub.status.busy":"2023-09-09T23:24:36.733904Z","iopub.execute_input":"2023-09-09T23:24:36.734307Z","iopub.status.idle":"2023-09-09T23:24:39.691089Z","shell.execute_reply.started":"2023-09-09T23:24:36.734265Z","shell.execute_reply":"2023-09-09T23:24:39.690353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. 視覚化\n\n人間の聴覚には、場所的（https://en.wikipedia.org/wiki/Place_theory_)（周波数ベース）と時間的（https://en.wikipedia.org/wiki/Temporal_theory_　）の2つの理論がある。音声認識では、スペクトログラム（周波数）(https://en.wikipedia.org/wiki/Spectrogram)を入力することと、より洗練された特徴MFCC（メル-周波数セプストラル係数）、PLPを入力することです。生の時間データを扱うことはほとんどありません。\n\nいくつかの録音を視覚化してみよう！\n\n1.1. 波形とスペクトログラム：\n\nファイルを選んで読み込む：\n\n# 1. Visualization \n<a id=\"visualization\"></a> \n\nThere are two theories of a human hearing - place ( https://en.wikipedia.org/wiki/Place_theory_(hearing) (frequency-based) and temporal (https://en.wikipedia.org/wiki/Temporal_theory_(hearing) )\nIn speech recognition, I see two main tendencies - to input [spectrogram](https://en.wikipedia.org/wiki/Spectrogram) (frequencies), and more sophisticated features MFCC - Mel-Frequency Cepstral Coefficients, PLP. You rarely work with raw, temporal data.\n\nLet's visualize some recordings!\n\n## 1.1. Wave and spectrogram:\n<a id=\"waveandspectrogram\"></a> \n\nChoose and read some file:","metadata":{"_uuid":"95fabaca63ab1a486bcc1f6824b26919ef325ff4","_cell_guid":"7f050711-6810-4aac-a306-da82fddb5579"}},{"cell_type":"code","source":"#train_audio_path = '../input/train/audio/'\ntrain_audio_path = '../input/speechcommandv02-cleaned/nyumaya_speech_command_V2_0/good_files/'\nfilename = '/yes/0a7c2a8d_nohash_0.wav'\nsample_rate, samples = wavfile.read(str(train_audio_path) + filename)","metadata":{"_uuid":"76266716e7df45a83073fb2964218c85b36d31cb","_cell_guid":"02126a6d-dd84-4f0a-88eb-ed9ff46a9bdf","execution":{"iopub.status.busy":"2023-09-09T23:24:39.692387Z","iopub.execute_input":"2023-09-09T23:24:39.692834Z","iopub.status.idle":"2023-09-09T23:24:39.708413Z","shell.execute_reply.started":"2023-09-09T23:24:39.692789Z","shell.execute_reply":"2023-09-09T23:24:39.707769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samples #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:39.709454Z","iopub.execute_input":"2023-09-09T23:24:39.709865Z","iopub.status.idle":"2023-09-09T23:24:39.716227Z","shell.execute_reply.started":"2023-09-09T23:24:39.709821Z","shell.execute_reply":"2023-09-09T23:24:39.715520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"与えられたデータセットは7zipで、解凍の仕方が分からなかったため、別のデータセットを使用した。\n以下のコードで解凍できたが、testはファイルが大きすぎ。また、trainデータの取り込みも難しいので諦めた。\n\n!pip install pyunpack\n!pip install patool\n\nimport os\nfrom pyunpack import Archive\nimport shutil\n\nif not os.path.exists('/kaggle/working/'):\n    os.makedirs('/kaggle/working/')\n\nArchive('/kaggle/input/tensorflow-speech-recognition-challenge/train.7z').extractall('/kaggle/working/')  \n\\# for dirname, _, filenames in os.walk('/kaggle/working/train/'):  \n\\#     for filename in filenames:  \n\\#         print(os.path.join(dirname, filename))  ","metadata":{}},{"cell_type":"markdown","source":"スペクトログラムを計算する関数を定義します。\n\nスペクトログラム値の対数を取っていることに注意してください。 それは私たちのプロットをより明確にするでしょう、そしてそれは人々の聴覚と厳密に関連しています。 対数への入力として 0 値がないことを確認する必要があります。\n\nDefine a function that calculates spectrogram.\n\nNote, that we are taking logarithm of spectrogram values. It will make our plot much more clear, moreover, it is strictly connected to the way people hear.\nWe need to assure that there are no 0 values as input to logarithm.","metadata":{"_uuid":"3bc26d76ea9f627c4d476ff8e9523f37d0668bbf","_cell_guid":"a7715152-3866-48dd-8bbb-31a72e9aa9bf"}},{"cell_type":"code","source":"def log_specgram(audio, sample_rate, window_size=20,\n                 step_size=10, eps=1e-10):\n    nperseg = int(round(window_size * sample_rate / 1e3))\n    noverlap = int(round(step_size * sample_rate / 1e3))\n    freqs, times, spec = signal.spectrogram(audio,\n                                    fs=sample_rate,\n                                    window='hann',\n                                    nperseg=nperseg,\n                                    noverlap=noverlap,\n                                    detrend=False)\n    return freqs, times, np.log(spec.T.astype(np.float32) + eps)","metadata":{"_uuid":"a3569f66d5bbbdcf338eaa121328a507f3a7b431","_cell_guid":"e464fe63-138e-4c66-a1f7-3ad3a81daa38","execution":{"iopub.status.busy":"2023-09-09T23:24:39.717444Z","iopub.execute_input":"2023-09-09T23:24:39.718147Z","iopub.status.idle":"2023-09-09T23:24:39.726628Z","shell.execute_reply.started":"2023-09-09T23:24:39.717974Z","shell.execute_reply":"2023-09-09T23:24:39.726046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ナイキスト定理によれば、周波数は (0、8000) の範囲内にあります。(https://en.wikipedia.org/wiki/Nyquist_rate)\nそれをプロットしてみましょう:\n\nナイキスト定理は、波形を正確に再構築するために信号の最大周波数成分の2倍以上の速さで信号をサンプルする必要があると定義します。 サンプリングレートの1/2の周波数を超える周波数成分が誤って低周波数成分と解釈される現象をエイリアスと呼びます。\n\nFrequencies are in range (0, 8000) according to [Nyquist theorem].\n\nLet's plot it:","metadata":{"_uuid":"4fd53946fd96b09765a267231ea5a66b313c2d4e","_cell_guid":"625dcb59-00ec-4b3f-97d5-f8adc12ac61a"}},{"cell_type":"code","source":"freqs, times, spectrogram = log_specgram(samples, sample_rate)\n\nfig = plt.figure(figsize=(14, 8))\nax1 = fig.add_subplot(211)\nax1.set_title('Raw wave of ' + filename)\nax1.set_ylabel('Amplitude')\nax1.plot(np.linspace(0, sample_rate/len(samples), sample_rate), samples)\n#連番や等差数列を生成するnumpy.linspace()。要素数を指定\n\nax2 = fig.add_subplot(212)\nax2.imshow(spectrogram.T, aspect='auto', origin='lower', \n           extent=[times.min(), times.max(), freqs.min(), freqs.max()])\n#imshow引数 (1/元の画像のアスペクト比)*やりたいアスペクト比をaspect引数に与える。\n#(アスペクト比は縦/横で考えることとする)\n#考え方としては、\n#画像を正方形にする(1/元の画像のアスペクト比で実現)\n#目的のアスペクト比をかける(やりたいアスペクト比をかけることで実現)\n#alpha引数から透明度を設定することができる\n#origin引数には'upper' or 'lower'を与えることが出来ます。デフォルトは'upper'です。\n#配列の[0,0]を'upper'の時は左上隅から配置し、'lower'の時は左下隅から配置してくれます。\n#extent引数を設定することで、画像が塗りつぶす範囲を明示することが出来ます。\n#引数への与え方は (left, right, bottom, top)です。\n#塗りつぶす範囲の決め方によってアスペクト比も変化するので、注意\n\nax2.set_yticks(freqs[::16])\nax2.set_xticks(times[::16])\nax2.set_title('Spectrogram of ' + filename)\nax2.set_ylabel('Freqs in Hz')\nax2.set_xlabel('Seconds')","metadata":{"_uuid":"ec1d065704c51f7f2b5b49f00da64809257815ab","_cell_guid":"4f77267a-1720-439b-9ef9-90e60f4446e1","execution":{"iopub.status.busy":"2023-09-09T23:24:39.727991Z","iopub.execute_input":"2023-09-09T23:24:39.728578Z","iopub.status.idle":"2023-09-09T23:24:40.219990Z","shell.execute_reply.started":"2023-09-09T23:24:39.728494Z","shell.execute_reply":"2023-09-09T23:24:40.219121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(freqs) #★\nprint(times) #★\nprint(spectrogram) #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:40.221490Z","iopub.execute_input":"2023-09-09T23:24:40.222068Z","iopub.status.idle":"2023-09-09T23:24:40.236943Z","shell.execute_reply.started":"2023-09-09T23:24:40.222012Z","shell.execute_reply":"2023-09-09T23:24:40.236139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"NumPyで連番や等差数列（等間隔の配列ndarray）を生成するにはnumpy.arange()かnumpy.linspace()を使う。arange()は間隔（公差）を指定、linspace()は要素数を指定という違いがあるので、目的によって使い分ける\n\n間隔（公差）を指定するnumpy.arange()の使い方\n要素数を指定するnumpy.linspace()の使い方\n基本的な使い方 stopを含むか指定: 引数endpoint 間隔を取得: 引数retstep","metadata":{}},{"cell_type":"markdown","source":"scipy.signal.spectrogram(x, fs=1.0, window=('tukey', 0.25), nperseg=None, noverlap=None, nfft=None, detrend='constant', return_onesided=True, scaling='density', axis=-1, mode='psd')source\nCompute a spectrogram with consecutive Fourier transforms.\n\nSpectrograms can be used as a way of visualizing the change of a nonstationary signal’s frequency content over time.\n\n連続したフーリエ変換を使用してスペクトログラムを計算します。\n\nスペクトログラムは、非定常信号の周波数成分の時間の経過に伴う変化を視覚化する方法として使用できます。\n\nパラメーター：\nxarray_like 測定値の時系列\n\nfsfloat、オプション\nx 時系列のサンプリング周波数。 デフォルトは 1.0 です。\n\nwindowstr または tuple または array_like、オプション\n使用する希望のウィンドウ。 window が文字列またはタプルの場合、それが get_window に渡されてウィンドウ値が生成されます。デフォルトでは DFT 偶数です。 ウィンドウと必要なパラメータのリストについては、get_window を参照してください。 window が array_like の場合、それはウィンドウとして直接使用され、その長さは nperseg でなければなりません。 デフォルトは、形状パラメータが 0.25 の Tukey ウィンドウです。\n\nnpersegint、オプション\n各セグメントの長さ。 デフォルトは None ですが、window が str または tuple の場合は 256 に設定され、window が array_like の場合はウィンドウの長さに設定されます。\n\nnoverlapint、オプション\nセグメント間で重なる点の数。 None の場合、noverlap = nperseg // 8. デフォルトは None です。\n\nnfftint、オプション\nゼロ埋め込み FFT が必要な場合に使用される FFT の長さ。 None の場合、FFT 長は nperseg です。 デフォルトは「なし」です。\n\ndetrendstr または関数または False、オプション\n各セグメントのトレンドを除去する方法を指定します。 detrend が文字列の場合、型引数として detrend 関数に渡されます。 関数の場合は、セグメントを取得し、傾向除去されたセグメントを返します。 detrend が False の場合、トレンド除去は行われません。 デフォルトは「定数」です。\n\nreturn_onesidebool、オプション\nTrue の場合、実データの片側スペクトルを返します。 False の場合、両側スペクトルを返します。 デフォルトは True ですが、複雑なデータの場合は常に両側スペクトルが返されます。\n\nスケーリング{ ‘密度’, ‘スペクトル’ }、オプション\nx が V で測定され、fs が Hz で測定される場合、Sxx の単位が V2/Hz であるパワー スペクトル密度 (「密度」) を計算するか、Sxx の単位が V2 であるパワー スペクトル (「スペクトル」) を計算するかを選択します。 デフォルトは「密度」です。\n\naxisint、オプション\nスペクトログラムが計算される軸。 デフォルトは最後の軸上です (つまり、axis=-1)。\n\nmodestr、オプション\nどのような種類の戻り値が予期されるかを定義します。 オプションは ['psd'、'complex'、'magnitude'、'angle'、'phase'] です。 「complex」は、パディングや境界拡張のない stft の出力と同等です。 「magnitude」は STFT の絶対的な大きさを返します。 「angle」と「phase」は、それぞれアンラップありとなしの STFT の複素角度を返します。\n\n戻り値：\nfndarray 太字 サンプル周波数の配列。\n\ntndarray\nセグメント時間の配列。\n\nSxxndarray\nxのスペクトログラム。 デフォルトでは、Sxx の最後の軸はセグメント時間に対応します。\n\nParameters:\nxarray_like\nTime series of measurement values\n\nfsfloat, optional\nSampling frequency of the x time series. Defaults to 1.0.\n\nwindowstr or tuple or array_like, optional\nDesired window to use. If window is a string or tuple, it is passed to get_window to generate the window values, which are DFT-even by default. See get_window for a list of windows and required parameters. If window is array_like it will be used directly as the window and its length must be nperseg. Defaults to a Tukey window with shape parameter of 0.25.\n\nnpersegint, optional\nLength of each segment. Defaults to None, but if window is str or tuple, is set to 256, and if window is array_like, is set to the length of the window.\n\nnoverlapint, optional\nNumber of points to overlap between segments. If None, noverlap = nperseg // 8. Defaults to None.\n\nnfftint, optional\nLength of the FFT used, if a zero padded FFT is desired. If None, the FFT length is nperseg. Defaults to None.\n\ndetrendstr or function or False, optional\nSpecifies how to detrend each segment. If detrend is a string, it is passed as the type argument to the detrend function. If it is a function, it takes a segment and returns a detrended segment. If detrend is False, no detrending is done. Defaults to ‘constant’.\n\nreturn_onesidedbool, optional\nIf True, return a one-sided spectrum for real data. If False return a two-sided spectrum. Defaults to True, but for complex data, a two-sided spectrum is always returned.\n\nscaling{ ‘density’, ‘spectrum’ }, optional\nSelects between computing the power spectral density (‘density’) where Sxx has units of V2/Hz and computing the power spectrum (‘spectrum’) where Sxx has units of V2, if x is measured in V and fs is measured in Hz. Defaults to ‘density’.\n\naxisint, optional\nAxis along which the spectrogram is computed; the default is over the last axis (i.e. axis=-1).\n\n*modestr, optional * Defines what kind of return values are expected. Options are [‘psd’, ‘complex’, ‘magnitude’, ‘angle’, ‘phase’]. ‘complex’ is equivalent to the output of stft with no padding or boundary extension. ‘magnitude’ returns the absolute magnitude of the STFT. ‘angle’ and ‘phase’ return the complex angle of the STFT, with and without unwrapping, respectively.\n\nReturns:\nfndarra\nArray of sample frequencies.\n\ntndarray\nArray of segment times.\n\nSxxndarray\nSpectrogram of x. By default, the last axis of Sxx corresponds to the segment times.","metadata":{}},{"cell_type":"markdown","source":"スペクトログラムを NN の入力特徴として使用する場合は、特徴を正規化することを忘れないでください。 (すべてのデータセットを正規化する必要があります。ここでは 1 つの例だけを示しますが、これでは良好な平均値と標準偏差が得られません。)\n\nIf we use spectrogram as an input features for NN, we have to remember to normalize features. (We need to normalize over all the dataset, here's example just for one, which doesn't give good *mean* and *std*!)","metadata":{"_uuid":"8f36fd74c9ad998d71b3a8838347b1bdbe8c82a7","_cell_guid":"013846a9-a929-45d9-97f5-98c59c6b2f23"}},{"cell_type":"code","source":"mean = np.mean(spectrogram, axis=0)\nstd = np.std(spectrogram, axis=0)\nspectrogram = (spectrogram - mean) / std","metadata":{"_uuid":"b3d09cf8bd1e91f54774f84dc952508d8a8a4eb8","_cell_guid":"9572b5e1-0b0f-42aa-934e-a1e313c21f46","execution":{"iopub.status.busy":"2023-09-09T23:24:40.238011Z","iopub.execute_input":"2023-09-09T23:24:40.238433Z","iopub.status.idle":"2023-09-09T23:24:40.242726Z","shell.execute_reply.started":"2023-09-09T23:24:40.238388Z","shell.execute_reply":"2023-09-09T23:24:40.242045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(mean) #★\nprint(std) #★\nprint(spectrogram) #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:40.243778Z","iopub.execute_input":"2023-09-09T23:24:40.244264Z","iopub.status.idle":"2023-09-09T23:24:40.261863Z","shell.execute_reply.started":"2023-09-09T23:24:40.244207Z","shell.execute_reply":"2023-09-09T23:24:40.261097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"指摘すべき興味深い事実があります。 各フレームには約 160 個の特徴があり、周波数は 0 ～ 8000 です。つまり、1 つの特徴は 50 Hz に対応します。 ただし、耳の周波数分解能は 1000 ～ 2000 Hz のオクターブ内で 3.6 Hz です。(https://en.wikipedia.org/wiki/Psychoacoustics) これは、人間は上記のようなスペクトログラムで表されるものよりもはるかに正確で、はるかに小さな詳細を聞き取ることができることを意味します。\n\nThere is an interesting fact to point out. We have ~160 features for each frame, frequencies are between 0 and 8000. It means, that one feature corresponds to 50 Hz. However, [frequency resolution of the ear is 3.6 Hz within the octave of 1000 – 2000 Hz](https://en.wikipedia.org/wiki/Psychoacoustics) It means, that people are far more precise and can hear much smaller details than those represented by spectrograms like above.","metadata":{"_uuid":"ee5c8c77122ece10687bacd92cd7ec01ab77af75","_cell_guid":"e3acd63a-4e3d-48fc-ba01-c28fedcd4496"}},{"cell_type":"markdown","source":"## 1.2. MFCC\n\nMFCC について詳しく知りたい場合は、この素晴らしいチュートリアルをご覧ください。 MFCC は、人間の聴覚特性を模倣する準備が十分に整っていることがわかります。[MFCC explained](http://practicalcryptography.com/miscellaneous/machine-learning/guide-mel-frequency-cepstral-coefficients-mfccs/)\n\nたとえば、librosa Python パッケージを使用して、メル パワー スペクトログラムと MFCC を計算できます。\n\n## 1.2. MFCC\n<a id=\"mfcc\"></a> \n\nIf you want to get to know some details about *MFCC* take a look at this great tutorial. [MFCC explained](http://practicalcryptography.com/miscellaneous/machine-learning/guide-mel-frequency-cepstral-coefficients-mfccs/) You can see, that it is well prepared to imitate human hearing properties.\n\nYou can calculate *Mel power spectrogram* and *MFCC* using for example *librosa* python package.\n","metadata":{"_uuid":"4eb99845d61397b9acb2488d34e2bafa7aad4cca","_cell_guid":"53904969-d453-4f0e-8e9b-6d932190bed1"}},{"cell_type":"code","source":"# From this tutorial\n# https://github.com/librosa/librosa/blob/master/examples/LibROSA%20demo.ipynb\nS = librosa.feature.melspectrogram(samples, sr=sample_rate, n_mels=128)\n#Google colabではこの部分がエラーになる。BINGによると、melspectrogram関数の位置引数の問題・・・\n\n# Convert to log scale (dB). We'll use the peak power (max) as reference.\nlog_S = librosa.power_to_db(S, ref=np.max)\n\n#librosa.power_to_db(S, *, ref=1.0, amin=1e-10, top_db=80.0)[source]\n#Convert a power spectrogram (amplitude squared) to decibel (dB) units\n#This computes the scaling 10 * log10(S / ref) in a numerically stable way.\n\nplt.figure(figsize=(12, 4))\nlibrosa.display.specshow(log_S, sr=sample_rate, x_axis='time', y_axis='mel')\nplt.title('Mel power spectrogram ')\nplt.colorbar(format='%+02.0f dB')\nplt.tight_layout()","metadata":{"_uuid":"4d996e6499140446685a3796418faa15a5f9d425","_cell_guid":"a6cb80ed-0e64-43b5-87ae-d33f3f844276","execution":{"iopub.status.busy":"2023-09-09T23:24:40.263298Z","iopub.execute_input":"2023-09-09T23:24:40.263852Z","iopub.status.idle":"2023-09-09T23:24:40.752062Z","shell.execute_reply.started":"2023-09-09T23:24:40.263794Z","shell.execute_reply":"2023-09-09T23:24:40.751177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mfcc = librosa.feature.mfcc(S=log_S, n_mfcc=13)\n\n# Let's pad on the first and second deltas while we're at it\ndelta2_mfcc = librosa.feature.delta(mfcc, order=2)\n\nplt.figure(figsize=(12, 4))\nlibrosa.display.specshow(delta2_mfcc)\nplt.ylabel('MFCC coeffs')\nplt.xlabel('Time')\nplt.title('MFCC')\nplt.colorbar()\nplt.tight_layout()","metadata":{"_uuid":"e1a21ec3fdbb30b360479d8886e3e496a5511ba4","_cell_guid":"38c436c0-9db0-48f8-a3d2-8ad660447bea","execution":{"iopub.status.busy":"2023-09-09T23:24:40.753429Z","iopub.execute_input":"2023-09-09T23:24:40.753914Z","iopub.status.idle":"2023-09-09T23:24:40.922306Z","shell.execute_reply.started":"2023-09-09T23:24:40.753842Z","shell.execute_reply":"2023-09-09T23:24:40.921426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"https://www.wizard-notes.com/entry/music-analysis/insts-timbre-with-mfcc  \n**MFCC**は音声認識や音楽ジャンル分類などで使われる特徴量であり、人間の聴覚特性を考慮した周波数スペクトルの概形（包絡線）を表しています。MFCCは楽器音に対しては音色に対応しており、音色が異なるとMFCCの形状は異なることが期待されます。\n\nMFCC だけを抽出したいときには、**librosa.feature.mfcc** が便利です。\nNumpy配列の時間信号を引数として渡すだけで、MFCCが算出できます。\nチューニングするパラメタとして重要なのが、MFCCの次数（特徴ベクトルの次元数）n_mfccです。n_mfccを大きい値にすると、メル周波数スペクトル包絡のより細かい成分まで考慮することができます。ただし、特徴ベクトルとして次元数が増えてしまうこともあり、分析や機械学習で使う場合は12～24くらいの次元数がよく使われます。","metadata":{}},{"cell_type":"markdown","source":"古典的ではあるが依然として最先端のシステムでは、スペクトログラムの代わりに MFCC または同様の機能がシステムへの入力として使用されます。\n\nただし、エンドツーエンド (多くの場合、ニューラル ネットワーク ベース) システムでは、最も一般的な入力特徴はおそらく生のスペクトログラム、またはメル パワー スペクトログラムです。 たとえば、MFCC は特徴を非相関化しますが、NN は相関のある特徴を適切に処理します。 また、メル フィルターを理解できれば、その使用法は賢明であると考えることができるでしょう。\n\nどちらを選ぶかはあなたの決断です！\n\nIn classical, but still state-of-the-art systems, *MFCC*  or similar features are taken as the input to the system instead of spectrograms.\n\nHowever, in end-to-end (often neural-network based) systems, the most common input features are probably raw spectrograms, or mel power spectrograms. For example *MFCC* decorrelates features, but NNs deal with correlated features well. Also, if you'll understand mel filters, you may consider their usage sensible.a\n\nIt is your decision which to choose!","metadata":{"_uuid":"ea023dd3edc2aea6a82b20eb4c53aac7f818390e","_cell_guid":"d1db710e-20d5-40ee-ac57-2e1e5de19c0e"}},{"cell_type":"markdown","source":"## 1.3. 3D のスペクトログラム¶\n\nところで、時代が変われば道具も変わります。 3D のスペクトログラムを見たことはありますか?\n\n## 1.3. Spectrogram in 3d\n<a id=\"3d\"></a> \n\nBy the way, times change, and the tools change. Have you ever seen spectrogram in 3d?","metadata":{"_uuid":"62f174047f35264a71dad240503d319a452271f9","_cell_guid":"156af766-21f8-4f93-ba1a-750157a11171"}},{"cell_type":"code","source":"data = [go.Surface(z=spectrogram.T)]\n#面グラフのデータを格納する Surfaceメソッド\n#plotlyで面グラフを描画するにはgraph_objectsに面グラフのデータを格納する必要があります。\n#面グラフのデータを作成するのに用いられるのがgraph_objectsのSurfaceメソッドです。\n\n #上のdataの意味や、下の表示のための引数が分からないが、まずは先へ。★★\n\nlayout = go.Layout(\n    title='Specgtrogram of \"yes\" in 3d',\n    scene = dict(\n    yaxis = dict(title='Frequencies', range=freqs), #Google colabではエラー発生\n    xaxis = dict(title='Time', range=times), #Google colabではエラー発生\n    zaxis = dict(title='Log amplitude'),\n    ),\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","metadata":{"_uuid":"cfb8b4827cecd2e01a1416176035488edd454fe4","_cell_guid":"51a5bc43-5216-4f13-8bc5-f84e005a01df","execution":{"iopub.status.busy":"2023-09-09T23:24:40.923605Z","iopub.execute_input":"2023-09-09T23:24:40.924122Z","iopub.status.idle":"2023-09-09T23:24:41.104012Z","shell.execute_reply.started":"2023-09-09T23:24:40.924056Z","shell.execute_reply":"2023-09-09T23:24:41.103216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spectrogram #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:41.105187Z","iopub.execute_input":"2023-09-09T23:24:41.105605Z","iopub.status.idle":"2023-09-09T23:24:41.113357Z","shell.execute_reply.started":"2023-09-09T23:24:41.105555Z","shell.execute_reply":"2023-09-09T23:24:41.112535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spectrogram.T #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:41.115035Z","iopub.execute_input":"2023-09-09T23:24:41.115670Z","iopub.status.idle":"2023-09-09T23:24:41.126376Z","shell.execute_reply.started":"2023-09-09T23:24:41.115609Z","shell.execute_reply":"2023-09-09T23:24:41.125660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"go.Surface(z=spectrogram.T) #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:41.127841Z","iopub.execute_input":"2023-09-09T23:24:41.128319Z","iopub.status.idle":"2023-09-09T23:24:41.137473Z","shell.execute_reply.started":"2023-09-09T23:24:41.128266Z","shell.execute_reply":"2023-09-09T23:24:41.136539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data #★data = [go.Surface(z=spectrogram.T)]","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:41.138520Z","iopub.execute_input":"2023-09-09T23:24:41.138875Z","iopub.status.idle":"2023-09-09T23:24:41.149681Z","shell.execute_reply.started":"2023-09-09T23:24:41.138834Z","shell.execute_reply":"2023-09-09T23:24:41.149085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"(軸範囲を適切な値に設定する方法がまだわかりません。また、上記の古典的なスペクトログラムのように拡張したいと思っています。)・・・同感★★\n\n(Don't know how to set axis ranges to proper values yet. I'd also like it to be streched like a classic spectrogram above..)","metadata":{"_uuid":"790e79ac09b6cea0b509965a86760bdcbe2671a7","_cell_guid":"7b8a7587-f41a-4021-9694-637f29e2c3f8"}},{"cell_type":"markdown","source":"## 1.4. Silence removal\n<a id=\"silenceremoval\"></a> \n\nLet's listen to that file","metadata":{"_uuid":"769e6738c4dae9923b9c0b0a99981bce8b443030","_cell_guid":"a2ad2019-f402-4226-9bad-65fb400aa8b1"}},{"cell_type":"code","source":"ipd.Audio(samples, rate=sample_rate)\n#IPython.display.Audio を使うと、\n#パスを指定すればそれを読み、データを指定すればそれを埋め込んでくれる。","metadata":{"_uuid":"ab0145dc0c8efdc08b4153b136c2b78634f6ed07","_cell_guid":"f49b916e-53a2-4dbe-bd03-3d8d93bf25a6","execution":{"iopub.status.busy":"2023-09-09T23:24:41.150651Z","iopub.execute_input":"2023-09-09T23:24:41.151025Z","iopub.status.idle":"2023-09-09T23:24:41.170201Z","shell.execute_reply.started":"2023-09-09T23:24:41.150984Z","shell.execute_reply":"2023-09-09T23:24:41.169651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ここでは、VAD (音声アクティビティ検出) が非常に役立つと考えています。 言葉は短いですが、その中には沈黙がたくさんあります。 適切な VAD を使用すると、トレーニングのサイズが大幅に削減され、トレーニング速度が大幅に加速されます。 ファイルの最初と最後から少し切り取ってみましょう。 そしてもう一度聞いてみましょう (上記のプロットに基づいて、4000 から 13000 までとります)。\n\nI consider that some *VAD* (Voice Activity Detection) will be really useful here. Although the words are short, there is a lot of silence in them. A decent *VAD* can reduce training size a lot, accelerating training speed significantly.\nLet's cut a bit of the file from the beginning and from the end. and listen to it again (based on a plot above, we take from 4000 to 13000):","metadata":{"_uuid":"9745fb19ce26c85c312a20e7fa19d98e672ceb64","_cell_guid":"4c23e8b3-0c8f-4eda-8f35-7486bdecfd9d"}},{"cell_type":"code","source":"samples_cut = samples[4000:13000]\nipd.Audio(samples_cut, rate=sample_rate)","metadata":{"_uuid":"539d123d84ac0181b820cca82c4098ab0ca54116","_cell_guid":"2c85e04d-cbd8-4702-bd50-7340497e800d","execution":{"iopub.status.busy":"2023-09-09T23:24:41.171091Z","iopub.execute_input":"2023-09-09T23:24:41.171433Z","iopub.status.idle":"2023-09-09T23:24:41.179257Z","shell.execute_reply.started":"2023-09-09T23:24:41.171391Z","shell.execute_reply":"2023-09-09T23:24:41.178581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(samples) #★\nprint(len(samples)) #★\nprint(samples.min()) #★\nprint(samples.max()) #★\nprint(samples_cut) #★\nprint(len(samples_cut)) #★\nprint(samples_cut.min()) #★\nprint(samples_cut.max()) #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:41.180253Z","iopub.execute_input":"2023-09-09T23:24:41.180670Z","iopub.status.idle":"2023-09-09T23:24:41.190948Z","shell.execute_reply.started":"2023-09-09T23:24:41.180614Z","shell.execute_reply":"2023-09-09T23:24:41.190209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"単語全体が聞こえることに同意できます。 すべてのファイルを手動で切り取り、単純なプロットに基づいてこれを行うことは不可能です。 ただし、たとえば webrtcvad パッケージを使用すると、優れた VAD を実現できます。\n\n'y' 'e' 's' グラフェムの推測された配置とともに、もう一度プロットしてみましょう\n\n書記素（しょきそ、英: **grapheme**）とは、書記言語において意味上の区別を可能にする最小の図形単位をいう。 口頭言語における音素に相当する。 字素（じそ）、文字素（もじそ）、図形素（ずけいそ）ともいう。 文字のほか、数字などの記号、あるいはそれらを構成する基本的要素を指す。\n\nWe can agree that the entire word can be heard. It is impossible to cut all the files manually and do this basing on the simple plot. But you can use for example *webrtcvad* package to have a good *VAD*.\n\nLet's plot it again, together with guessed alignment of* 'y' 'e' 's'* graphems","metadata":{"_uuid":"e9ceecbabecc11f789b3de382ee4c909186e6d22","_cell_guid":"45898236-4528-4e21-86dd-55abcf4f639f"}},{"cell_type":"code","source":"freqs, times, spectrogram_cut = log_specgram(samples_cut, sample_rate)\n\nfig = plt.figure(figsize=(14, 8))\nax1 = fig.add_subplot(211)\nax1.set_title('Raw wave of ' + filename)\nax1.set_ylabel('Amplitude')\nax1.plot(samples_cut)\n\nax2 = fig.add_subplot(212)\nax2.set_title('Spectrogram of ' + filename)\nax2.set_ylabel('Frequencies * 0.1')\nax2.set_xlabel('Samples')\nax2.imshow(spectrogram_cut.T, aspect='auto', origin='lower', \n           extent=[times.min(), times.max(), freqs.min(), freqs.max()])\nax2.set_yticks(freqs[::16])\nax2.set_xticks(times[::16])\nax2.text(0.06, 1000, 'Y', fontsize=18)\nax2.text(0.17, 1000, 'E', fontsize=18)\nax2.text(0.36, 1000, 'S', fontsize=18)\n\nxcoords = [0.025, 0.11, 0.23, 0.49]\nfor xc in xcoords:\n    ax1.axvline(x=xc*16000, c='r')\n    ax2.axvline(x=xc, c='r')\n\n#axvline – 垂直な線分を1本引く","metadata":{"_uuid":"6831fa9311397dc8bca4192f657767d36c5c1a38","_cell_guid":"038fe488-4f25-42bd-af11-108f5ecbb1e7","execution":{"iopub.status.busy":"2023-09-09T23:24:41.192310Z","iopub.execute_input":"2023-09-09T23:24:41.192790Z","iopub.status.idle":"2023-09-09T23:24:41.621375Z","shell.execute_reply.started":"2023-09-09T23:24:41.192724Z","shell.execute_reply":"2023-09-09T23:24:41.620548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(freqs) #★\nprint(times) #★\nprint(spectrogram_cut) #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:41.622453Z","iopub.execute_input":"2023-09-09T23:24:41.622883Z","iopub.status.idle":"2023-09-09T23:24:41.633260Z","shell.execute_reply.started":"2023-09-09T23:24:41.622820Z","shell.execute_reply":"2023-09-09T23:24:41.632391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.5. リサンプリング - 次元削減\n<a id=\"resampl\"></a> \n\nデータの次元を削減するもう 1 つの方法は、録音をリサンプリングすることです。\n\n録音は 16k の周波数でサンプリングされているため、あまり自然に聞こえないことがわかりますが、通常はそれよりもはるかに多くの周波数が聞こえます。 ただし、音声に関連する周波数のほとんどは、より小さな帯域で表示されます。[the most speech related frequencies are presented in smaller band](https://en.wikipedia.org/wiki/Voice_frequency) GSM 信号が 8000 Hz でサンプリングされている電話で他の人が話していることを理解できるのはこのためです。\n\n要約すると、データセットを 8k にリサンプリングできます。 重要ではない情報の一部を破棄し、データのサイズを削減します。\n\nこれは競争であるため、リスクが伴う可能性があることを覚えておく必要があり、パフォーマンスの非常に小さな差が勝利する場合もあるため、何も失いたくないのです。 一方、最初の実験は、トレーニング サイズが小さいほど、はるかに高速に実行できます。\n\nFFT (高速フーリエ変換) を計算する必要があります。 意味：\n\n## 1.5. Resampling - dimensionality reduction\n<a id=\"resampl\"></a> \n\nAnother way to reduce the dimensionality of our data is to resample recordings.\n\nYou can hear that the recording don't sound very natural, because they are sampled with 16k frequency, and we usually hear much more. However, [the most speech related frequencies are presented in smaller band](https://en.wikipedia.org/wiki/Voice_frequency). That's why you can still understand another person talking to the telephone, where GSM signal is sampled to 8000 Hz.\n\nSummarizing, we could resample our dataset to 8k. We will discard some information that shouldn't be important, and we'll reduce size of the data.\n\nWe have to remember that it can be risky, because this is a competition, and sometimes very small difference in performance wins, so we don't want to lost anything. On the other hand, first experiments can be done much faster with smaller training size.\n\nWe'll need to calculate FFT (Fast Fourier Transform). Definition:\n","metadata":{"_uuid":"e8f5fa497bbd2b3f5e7dbb9fa20d59d9773309a1","_cell_guid":"f081f185-336a-429d-ba71-c0d2337c35ae"}},{"cell_type":"code","source":"def custom_fft(y, fs): #FFT (高速フーリエ変換) \n    T = 1.0 / fs\n    N = y.shape[0]\n    yf = fft(y)\n    xf = np.linspace(0.0, 1.0/(2.0*T), N//2)\n    vals = 2.0/N * np.abs(yf[0:N//2])  # FFT is simmetrical, so we take just the first half\n    # FFT is also complex, to we take just the real part (abs)\n    return xf, vals","metadata":{"_uuid":"213de6a783443d118c3509acc26f9f4bd0319d85","_cell_guid":"86dedd69-e084-403e-9c02-5370018acf1c","execution":{"iopub.status.busy":"2023-09-09T23:24:41.634308Z","iopub.execute_input":"2023-09-09T23:24:41.634716Z","iopub.status.idle":"2023-09-09T23:24:41.641803Z","shell.execute_reply.started":"2023-09-09T23:24:41.634655Z","shell.execute_reply":"2023-09-09T23:24:41.641107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"録音を読んでリサンプリングして聞いてみましょう。 FFT を比較することもできます。元の信号には 4000 Hz を超える情報がほとんどないことに注意してください。\n\nLet's read some recording, resample it, and listen. We can also compare FFT, Notice, that there is almost no information above 4000 Hz in original signal.","metadata":{"_uuid":"665e57b4652493e6d3b61ba2b7e70967170e7900","_cell_guid":"0fc3b446-d19e-4cd2-b1d6-3cf58ff332bf"}},{"cell_type":"code","source":"filename = '/happy/0b09edd3_nohash_0.wav'\nnew_sample_rate = 8000\n\nsample_rate, samples = wavfile.read(str(train_audio_path) + filename)\nresampled = signal.resample(samples, int(new_sample_rate/sample_rate * samples.shape[0]))","metadata":{"_uuid":"b8fdb36dc4fce089ea5a3c3dcc27f65625232e34","_cell_guid":"919e85ca-7769-4214-a1d7-5eaa74a32b19","execution":{"iopub.status.busy":"2023-09-09T23:24:41.643012Z","iopub.execute_input":"2023-09-09T23:24:41.643434Z","iopub.status.idle":"2023-09-09T23:24:41.659487Z","shell.execute_reply.started":"2023-09-09T23:24:41.643377Z","shell.execute_reply":"2023-09-09T23:24:41.658485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ipd.Audio(samples, rate=sample_rate)","metadata":{"_uuid":"afa8138a2ae7888ade44713fb5f8451f9c9e7f02","_cell_guid":"13f397f1-cd5d-4f0f-846a-0edd9f58bcff","execution":{"iopub.status.busy":"2023-09-09T23:24:41.660559Z","iopub.execute_input":"2023-09-09T23:24:41.660955Z","iopub.status.idle":"2023-09-09T23:24:41.672313Z","shell.execute_reply.started":"2023-09-09T23:24:41.660891Z","shell.execute_reply":"2023-09-09T23:24:41.671522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ipd.Audio(resampled, rate=new_sample_rate)","metadata":{"_uuid":"3f600c9414ab5cef205c814ba16a356d4121790b","_cell_guid":"5ab11b21-9528-47fa-8ff0-244b1d0c94b3","execution":{"iopub.status.busy":"2023-09-09T23:24:41.673399Z","iopub.execute_input":"2023-09-09T23:24:41.673770Z","iopub.status.idle":"2023-09-09T23:24:41.687319Z","shell.execute_reply.started":"2023-09-09T23:24:41.673728Z","shell.execute_reply":"2023-09-09T23:24:41.686512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Almost no difference!","metadata":{"_uuid":"96380594085d818693b959307d371e95f727f03b","_cell_guid":"37da8174-e6aa-463d-bef7-c8b20c6ca513"}},{"cell_type":"code","source":"xf, vals = custom_fft(samples, sample_rate)\nplt.figure(figsize=(12, 4))\nplt.title('FFT of recording sampled with ' + str(sample_rate) + ' Hz')\nplt.plot(xf, vals)\nplt.xlabel('Frequency')\nplt.grid()\nplt.show()","metadata":{"_uuid":"4448038dfa22ec582cde229346cb1ba309c76b9f","_cell_guid":"baed6102-3c75-4f16-85d7-723d8a084b9a","execution":{"iopub.status.busy":"2023-09-09T23:24:41.688240Z","iopub.execute_input":"2023-09-09T23:24:41.688590Z","iopub.status.idle":"2023-09-09T23:24:41.865760Z","shell.execute_reply.started":"2023-09-09T23:24:41.688550Z","shell.execute_reply":"2023-09-09T23:24:41.865040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"xf, vals = custom_fft(resampled, new_sample_rate)\nplt.figure(figsize=(12, 4))\nplt.title('FFT of recording sampled with ' + str(new_sample_rate) + ' Hz')\nplt.plot(xf, vals)\nplt.xlabel('Frequency')\nplt.grid()\nplt.show()","metadata":{"_uuid":"88953237ea59d13e9647813bef06a911f06f0e61","_cell_guid":"3cc1a49a-4cd4-49ed-83c8-f2437062f8be","execution":{"iopub.status.busy":"2023-09-09T23:24:41.866724Z","iopub.execute_input":"2023-09-09T23:24:41.867167Z","iopub.status.idle":"2023-09-09T23:24:42.044238Z","shell.execute_reply.started":"2023-09-09T23:24:41.867103Z","shell.execute_reply":"2023-09-09T23:24:42.043683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"このようにして、データセットのサイズを 2 倍削減しました。\n\nThis is how we reduced dataset size twice!","metadata":{"_uuid":"152c1b14d7a7b57d7ab4fb0bd52e38564406cb92","_cell_guid":"592ffc6a-edda-4b08-9419-d3462599da5c"}},{"cell_type":"markdown","source":"## 1.6. 特徴抽出手順\n<a id=\"featuresextractionsteps\"></a> \n\n私は次のような特徴抽出アルゴリズムを提案します。\n\n1. リサンプリング\n2. VAD\n3. 信号の長さを等しくするために 0 をパディングする可能性があります\n4. ログ スペクトログラム (または MFCC、または PLP)\n5. 平均値と標準値による正規化機能\n6. 一定数のフレームを積み重ねて時間情報を取得する\n\nノートではできないのが残念です。 ゼロから何かを書くのはあまり意味がなく、すべてが準備完了ですが、パッケージではそれをカーネルにインポートできません。\n\n## 1.6. Features extraction steps\n<a id=\"featuresextractionsteps\"></a> \n\nI would propose the feature extraction algorithm like that:\n1. Resampling\n2. *VAD*\n3. Maybe padding with 0 to make signals be equal length\n4. Log spectrogram (or *MFCC*, or *PLP*)\n5. Features normalization with *mean* and *std*\n6. Stacking of a given number of frames to get temporal information\n\nIt's a pity it can't be done in notebook. It has not much sense to write things from zero, and everything is ready to take, but in packages, that can not be imported in Kernels.","metadata":{"_uuid":"57fe8c6a25753e2eb46285bc8d725d20182c1421","_cell_guid":"f98fe35d-2d56-4153-b054-0882bd2e58ce"}},{"cell_type":"markdown","source":"# 2. データセットの調査\n<a id=\"investigations\"></a> \n\nデータセットの通常の調査。\n\n## 2.1. レコード数\n<a id=\"numberoffiles\"></a> \n\n# 2. Dataset investigation\n<a id=\"investigations\"></a> \n\nSome usuall investgation of dataset.\n\n## 2.1. Number of records\n<a id=\"numberoffiles\"></a> \n","metadata":{"_uuid":"caf345ca07983f1e1d4f8a05f6f74859554289db","_cell_guid":"3d36bac6-eb6f-4a53-b148-805493e39052"}},{"cell_type":"code","source":"dirs = [f for f in os.listdir(train_audio_path) if isdir(join(train_audio_path, f))]\ndirs.sort()\nprint('Number of labels: ' + str(len(dirs)))","metadata":{"_uuid":"59826a3eb0f60439d5beee06781193bc67cc53f7","_cell_guid":"3c24fbdd-e50e-47a1-8c44-1894bec7f043","execution":{"iopub.status.busy":"2023-09-09T23:24:42.045127Z","iopub.execute_input":"2023-09-09T23:24:42.045469Z","iopub.status.idle":"2023-09-09T23:24:42.055153Z","shell.execute_reply.started":"2023-09-09T23:24:42.045428Z","shell.execute_reply":"2023-09-09T23:24:42.054199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dirs #★","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:42.057655Z","iopub.execute_input":"2023-09-09T23:24:42.058135Z","iopub.status.idle":"2023-09-09T23:24:42.063339Z","shell.execute_reply.started":"2023-09-09T23:24:42.058076Z","shell.execute_reply":"2023-09-09T23:24:42.062457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate\nnumber_of_recordings = []\nfor direct in dirs:\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    number_of_recordings.append(len(waves))\n\n# Plot\ndata = [go.Histogram(x=dirs, y=number_of_recordings)]\ntrace = go.Bar(\n    x=dirs,\n    y=number_of_recordings,\n    marker=dict(color = number_of_recordings, colorscale='Viridius', showscale=True\n    ), # Google colabではエラー発生\n)\nlayout = go.Layout(\n    title='Number of recordings in given label',\n    xaxis = dict(title='Words'),\n    yaxis = dict(title='Number of recordings')\n)\npy.iplot(go.Figure(data=[trace], layout=layout))\n","metadata":{"_uuid":"ea30edeaf3d8020bf55ee2a57af230bded9732e2","_cell_guid":"a6b82ced-df8c-4c7a-8d4c-ed32bf9f60f6","execution":{"iopub.status.busy":"2023-09-09T23:24:42.064655Z","iopub.execute_input":"2023-09-09T23:24:42.065221Z","iopub.status.idle":"2023-09-09T23:24:45.157749Z","shell.execute_reply.started":"2023-09-09T23:24:42.065170Z","shell.execute_reply":"2023-09-09T23:24:45.157125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"データセットは、background_noise を除いてバランスが取れていますが、それは別のものです。\n\nDataset is balanced except of background_noise, but that's the different thing.","metadata":{"_uuid":"928d2933df5e0a37c7dc40f2ec50b9d10423d533","_cell_guid":"c4b1d377-d84e-4601-97fa-2c989023c400"}},{"cell_type":"markdown","source":"## 2.2. Deeper into recordings\n<a id=\"deeper\"></a> ","metadata":{"_uuid":"0847bb69f23ce4c8153f869222c03731fd62a22e","_cell_guid":"b9b6b43c-96fa-489d-8a49-b77dabd41705"}},{"cell_type":"markdown","source":"とても重要な事実があります。 録音は非常に異なるソースから来ています。 私の知る限り、それらの一部はモバイル GSM チャネルから送信されている可能性があります。\n\nそれにもかかわらず、**トレーニング セットとテスト セットの両方で 1 人の話者が出現しないようにデータセットを分割することが非常に重要です。**\nこの 2 つの例を見て聞いてください。\n\nThere's a very important fact. Recordings come from very different sources. As far as I can tell, some of them can come from mobile GSM channel.\n\nNevertheless,** it is extremely important to split the dataset in a way that one speaker doesn't occur in both train and test sets.**\nJust take a look and listen to this two examlpes:","metadata":{"_uuid":"d4b8b90afc03493f93babb7b9d401ebd0caa1c18","_cell_guid":"766c2e18-2aff-43c1-9672-bd51d4348867"}},{"cell_type":"code","source":"filenames = ['on/004ae714_nohash_0.wav', 'on/0137b3f4_nohash_0.wav']\nfor filename in filenames:\n    sample_rate, samples = wavfile.read(str(train_audio_path) + filename)\n    xf, vals = custom_fft(samples, sample_rate)\n    plt.figure(figsize=(12, 4))\n    plt.title('FFT of speaker ' + filename[4:11])\n    plt.plot(xf, vals)\n    plt.xlabel('Frequency')\n    plt.grid()\n    plt.show()","metadata":{"_uuid":"de451cfb747953d60d9b6982c046b417fdc1ab9b","_cell_guid":"d29ea00a-54eb-474a-8770-fbd269bbd21e","execution":{"iopub.status.busy":"2023-09-09T23:24:45.158829Z","iopub.execute_input":"2023-09-09T23:24:45.159148Z","iopub.status.idle":"2023-09-09T23:24:45.496890Z","shell.execute_reply.started":"2023-09-09T23:24:45.159088Z","shell.execute_reply":"2023-09-09T23:24:45.496095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even better to listen:","metadata":{"_uuid":"ae779083e33d29a22bcf972542d9911a7b9d64de","_cell_guid":"508c79b0-6beb-43cb-9092-9e0c5719e12e"}},{"cell_type":"code","source":"print('Speaker ' + filenames[0][4:11])\nipd.Audio(join(train_audio_path, filenames[0]))","metadata":{"_uuid":"4e2d4b6d2c6e6806a2b0b0d4554e49b080003a62","_cell_guid":"63e5851f-b462-4f9c-8e17-a7e8e301d616","execution":{"iopub.status.busy":"2023-09-09T23:24:45.498156Z","iopub.execute_input":"2023-09-09T23:24:45.498643Z","iopub.status.idle":"2023-09-09T23:24:45.509352Z","shell.execute_reply.started":"2023-09-09T23:24:45.498586Z","shell.execute_reply":"2023-09-09T23:24:45.508622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Speaker ' + filenames[1][4:11])\nipd.Audio(join(train_audio_path, filenames[1]))","metadata":{"_uuid":"c335e31977a721831149579da9105cc2b664b40d","_cell_guid":"fd2eb518-df30-44a0-89b3-9c4611beb50d","execution":{"iopub.status.busy":"2023-09-09T23:24:45.510303Z","iopub.execute_input":"2023-09-09T23:24:45.510881Z","iopub.status.idle":"2023-09-09T23:24:45.518007Z","shell.execute_reply.started":"2023-09-09T23:24:45.510818Z","shell.execute_reply":"2023-09-09T23:24:45.517423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"奇妙な無音 (何らかの圧縮?) が含まれる録音もあります。\n\nThere are also recordings with some weird silence (some compression?):\n","metadata":{"_uuid":"63555f55e906409e008d9c0a988a03b26cbb8983","_cell_guid":"7c2f7df0-d062-46d6-9507-5ab548db14bb"}},{"cell_type":"code","source":"filename = '/yes/01bb6a2a_nohash_1.wav'\nsample_rate, samples = wavfile.read(str(train_audio_path) + filename)\nfreqs, times, spectrogram = log_specgram(samples, sample_rate)\n\nplt.figure(figsize=(10, 7))\nplt.title('Spectrogram of ' + filename)\nplt.ylabel('Freqs')\nplt.xlabel('Time')\nplt.imshow(spectrogram.T, aspect='auto', origin='lower', \n           extent=[times.min(), times.max(), freqs.min(), freqs.max()])\nplt.yticks(freqs[::16])\nplt.xticks(times[::16])\nplt.show()","metadata":{"_uuid":"ba2e76ebe90eb9c4fed6617d1948e0c71d88c54e","_cell_guid":"bc6075ee-fc43-4d04-b400-64896ac8450d","execution":{"iopub.status.busy":"2023-09-09T23:24:45.519048Z","iopub.execute_input":"2023-09-09T23:24:45.519433Z","iopub.status.idle":"2023-09-09T23:24:45.816286Z","shell.execute_reply.started":"2023-09-09T23:24:45.519374Z","shell.execute_reply":"2023-09-09T23:24:45.815620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filename = '/yes/01bb6a2a_nohash_1.wav' #追加\nnew_sample_rate = 8000\n\nsample_rate, samples = wavfile.read(str(train_audio_path) + filename)\nresampled = signal.resample(samples, int(new_sample_rate/sample_rate * samples.shape[0]))\n\nipd.Audio(samples, rate=sample_rate)","metadata":{"execution":{"iopub.status.busy":"2023-09-09T23:24:45.817309Z","iopub.execute_input":"2023-09-09T23:24:45.817695Z","iopub.status.idle":"2023-09-09T23:24:45.831177Z","shell.execute_reply.started":"2023-09-09T23:24:45.817640Z","shell.execute_reply":"2023-09-09T23:24:45.830559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"つまり、非常に特殊な音響環境への過剰適合を防ぐ必要があります。\n\nIt means, that we have to prevent overfitting to the very specific acoustical environments.\n","metadata":{"_uuid":"cbf303188e33f325691aa9447f18586788bf110a","_cell_guid":"fb57dc50-0510-4dcf-ba07-7999daf7349e"}},{"cell_type":"markdown","source":"## 2.3. 録音の長さ¶\n\nすべてのファイルの長さが 1 秒かどうかを確認します。\n\n## 2.3. Recordings length\n<a id=\"len\"></a> \n\nFind if all the files have 1 second duration:","metadata":{"_uuid":"85c9f5b8e2dac9bf2aa2edf0ebebc0ae53ff6533","_cell_guid":"ae626131-069f-4243-b7fc-c0014b11e2d8"}},{"cell_type":"code","source":"num_of_shorter = 0\nfor direct in dirs:\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    for wav in waves:\n        sample_rate, samples = wavfile.read(train_audio_path + direct + '/' + wav)\n        if samples.shape[0] < sample_rate:\n            num_of_shorter += 1\nprint('Number of recordings shorter than 1 second: ' + str(num_of_shorter))","metadata":{"_uuid":"16a2e2c908235a99f64024abab272c65d3d99c65","_cell_guid":"23be7e5e-e4b4-40a0-b9a3-4bc850571a28","execution":{"iopub.status.busy":"2023-09-09T23:24:45.832452Z","iopub.execute_input":"2023-09-09T23:24:45.832873Z","iopub.status.idle":"2023-09-09T23:30:07.519773Z","shell.execute_reply.started":"2023-09-09T23:24:45.832829Z","shell.execute_reply":"2023-09-09T23:30:07.518802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"それは驚くべきことであり、それらはたくさんあります。 ゼロでPad(埋める?)ができます。\n\nThat's suprising, and there is a lot of them. We can pad them with zeros.","metadata":{"_uuid":"c035e161e9c9622fa96f9589ffbfe826e01c5658","_cell_guid":"57b26071-e603-4ee7-b1d5-5bd2bd46e438"}},{"cell_type":"markdown","source":"## 2.4. Mean spectrograms and FFT\n<a id=\"meanspectrogramsandfft\"></a> ","metadata":{"_uuid":"96655097f9173d34d529a0626446194473cebf69","_cell_guid":"771ff180-73dd-4be8-aa56-ccecb3586416"}},{"cell_type":"markdown","source":"すべての単語の平均FFTをプロットしましょう\n\nLet's plot mean FFT for every word","metadata":{"_uuid":"a3f64232afa284102e8ffdcdbe0db509f4a78a7e","_cell_guid":"a621a03b-4812-4d7c-93dd-6b0bf7f10572"}},{"cell_type":"code","source":"to_keep = 'yes no up down left right on off stop go'.split()\ndirs = [d for d in dirs if d in to_keep]\n\nprint(dirs)\n\nfor direct in dirs:\n    vals_all = []\n    spec_all = []\n\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    for wav in waves:\n        sample_rate, samples = wavfile.read(train_audio_path + direct + '/' + wav)\n        if samples.shape[0] != 16000:\n            continue\n        xf, vals = custom_fft(samples, 16000)\n        vals_all.append(vals)\n        freqs, times, spec = log_specgram(samples, 16000)\n        spec_all.append(spec)\n\n    plt.figure(figsize=(14, 4))\n    plt.subplot(121)\n    plt.title('Mean fft of ' + direct)\n    plt.plot(np.mean(np.array(vals_all), axis=0))\n    plt.grid()\n    plt.subplot(122)\n    plt.title('Mean specgram of ' + direct)\n    plt.imshow(np.mean(np.array(spec_all), axis=0).T, aspect='auto', origin='lower', \n               extent=[times.min(), times.max(), freqs.min(), freqs.max()])\n    plt.yticks(freqs[::16])\n    plt.xticks(times[::16])\n    plt.show()","metadata":{"_uuid":"e8fc86e267bb48d3ae35f8e2d85db070f28889c9","_cell_guid":"dc065096-6888-4a38-9c66-0729bfa858f6","execution":{"iopub.status.busy":"2023-09-09T23:30:07.521223Z","iopub.execute_input":"2023-09-09T23:30:07.521757Z","iopub.status.idle":"2023-09-09T23:31:33.692476Z","shell.execute_reply.started":"2023-09-09T23:30:07.521693Z","shell.execute_reply":"2023-09-09T23:31:33.691411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.5. ガウス混合モデリング¶\n\n*平均 FFT が単語ごとに異なって見える* ことがわかります。 **ガウス分布の混合を使用して各 FFT をモデル化** できます。 ただし、ストップとアップなど、FFT ではほとんど同じに見えるものもあります。しかし、待ってください。*スペクトログラムを見ると、まだ区別できます。 停止の開始時に高周波が低周波よりも早く* なります (おそらく s)。\n\nだからこそ **時間的な要素も必要** なのです。 *Kaldi ライブラリ* があり、**GMM で単語 (または単語の小さな部分) をモデル化し、隠れマルコフ モデルで時間的依存関係をモデル化** できます。\n\n**単語に単純な GMM を使用して、何をモデル化できるか、単語を区別するのがどれほど難しいかを確認** できます。 そのために Scikit-learn を使用することもできますが、これは簡単ではなく、非常に長く続くため、この考えは今のところ放棄します。\n\n## 2.5. Gaussian Mixtures modeling\n<a id=\"gmms\"></a> \n\nWe can see that mean FFT looks different for every word. We could model each FFT with a mixture of Gaussian distributions. Some of them however, look almost identical on FFT, like *stop* and *up*... But wait, they are still distinguishable when we look at spectrograms! High frequencies are earlier than low at the beginning of *stop* (probably *s*).\n\nThat's why temporal component is also necessary. There is a [Kaldi](http://kaldi-asr.org/) library, that can model words (or smaller parts of words) with GMMs and model temporal dependencies with [Hidden Markov Models](https://github.com/danijel3/ASRDemos/blob/master/notebooks/HMM_FST.ipynb).\n\nWe could use simple GMMs for words to check what can we model and how hard it is to distinguish the words. We can use [Scikit-learn](http://scikit-learn.org/) for that, however it is not straightforward and lasts very long here, so I abandon this idea for now.","metadata":{"_uuid":"931aceb2fb6ac23defc699b3d423b510171b1626","_cell_guid":"1089a473-7fe5-40a0-9e43-8f3cbf1901c3"}},{"cell_type":"markdown","source":"## 2.6. 2.6. 単語全体の周波数成分\n\n## Frequency components across the words\n<a id=\"components\"></a> \n","metadata":{"_uuid":"0c89774ecfd33f29c10cba58f1fb1c12647b0928","_cell_guid":"341fd72e-750e-4d78-a7b7-1d347bcad4e8"}},{"cell_type":"code","source":"def violinplot_frequency(dirs, freq_ind):\n    \"\"\" Plot violinplots for given words (waves in dirs) and frequency freq_ind\n    from all frequencies freqs.\"\"\"\n\n    spec_all = []  # Contain spectrograms\n    ind = 0\n    for direct in dirs:\n        spec_all.append([])\n\n        waves = [f for f in os.listdir(join(train_audio_path, direct)) if\n                 f.endswith('.wav')]\n                #特定の文字列で終わる（後方一致）: str.endswith()を使うと、\n                #要素が特定の文字列で終わるとTrueとなるpandas.Seriesを取得できる。\n        for wav in waves[:100]:\n            sample_rate, samples = wavfile.read(\n                train_audio_path + direct + '/' + wav)\n            freqs, times, spec = log_specgram(samples, sample_rate)\n            spec_all[ind].extend(spec[:, freq_ind])\n                #リストとリストを結合（連結）: extend(), +演算子\n                #リストのextend()メソッドで、リストに別のリストやタプルを結合できる。\n                #すべての要素が元のリストの末尾に追加される。\n        ind += 1\n\n    # Different lengths = different num of frames. Make number equal\n    minimum = min([len(spec) for spec in spec_all])\n    spec_all = np.array([spec[:minimum] for spec in spec_all])\n\n    plt.figure(figsize=(13,7))\n    plt.title('Frequency ' + str(freqs[freq_ind]) + ' Hz')\n    plt.ylabel('Amount of frequency in a word')\n    plt.xlabel('Words')\n    sns.violinplot(data=pd.DataFrame(spec_all.T, columns=dirs))\n    plt.show()\n    \n#データの分布密度を視覚化 violinplot\n#バイオリンプロットは、データの分布の密度を確認できるグラフとなっている。\n#matplotlib ライブラリーの violinplot メソッドを使って描くことができる。\n#入力データとして、リストの形式で与える。","metadata":{"_uuid":"0c3e4ecb1eb87bd8c86547df68db35ef9a84a2ce","_cell_guid":"7a883de7-1fa7-428d-bd92-0fb0d3f75aed","execution":{"iopub.status.busy":"2023-09-09T23:31:33.693915Z","iopub.execute_input":"2023-09-09T23:31:33.694285Z","iopub.status.idle":"2023-09-09T23:31:33.713825Z","shell.execute_reply.started":"2023-09-09T23:31:33.694222Z","shell.execute_reply":"2023-09-09T23:31:33.713207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"violinplot_frequency(dirs, 20)","metadata":{"_uuid":"8932034e13e12d7ff50bc6663c03153ae1052378","_cell_guid":"08aae780-53a3-4047-b548-46b8893aaed9","execution":{"iopub.status.busy":"2023-09-09T23:31:33.714794Z","iopub.execute_input":"2023-09-09T23:31:33.715189Z","iopub.status.idle":"2023-09-09T23:31:36.208885Z","shell.execute_reply.started":"2023-09-09T23:31:33.715129Z","shell.execute_reply":"2023-09-09T23:31:36.208301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"violinplot_frequency(dirs, 50)","metadata":{"_uuid":"30e1bf01aa6431fd883530beabb4da5d9c8518dc","_cell_guid":"ea7ea51e-3a38-46db-891d-4c3aef8fd810","execution":{"iopub.status.busy":"2023-09-09T23:31:36.209791Z","iopub.execute_input":"2023-09-09T23:31:36.210121Z","iopub.status.idle":"2023-09-09T23:31:38.527082Z","shell.execute_reply.started":"2023-09-09T23:31:36.210082Z","shell.execute_reply":"2023-09-09T23:31:38.526408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"violinplot_frequency(dirs, 120)","metadata":{"_uuid":"fb2a917be56bc19c256ad52e257122a7f1b5dd8b","_cell_guid":"a84e722a-3aaa-42dd-868b-6eb629abf40d","execution":{"iopub.status.busy":"2023-09-09T23:31:38.527971Z","iopub.execute_input":"2023-09-09T23:31:38.528283Z","iopub.status.idle":"2023-09-09T23:31:40.828344Z","shell.execute_reply.started":"2023-09-09T23:31:38.528245Z","shell.execute_reply":"2023-09-09T23:31:40.827765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.7. 異常検出¶\n\n何らかの形で*他の録音より際立った録音がないかどうかを確認する必要*があります。 *データセットの次元を下げて、異常がないか対話的にチェック*できます。 **次元削減には PCA を使用します。**\n\n## 2.7. Anomaly detection\n<a id=\"anomaly\"></a> \n\nWe should check if there are any recordings that somehow stand out from the rest. We can lower the dimensionality of the dataset and interactively check for any anomaly.\nWe'll use PCA for dimensionality reduction:","metadata":{"_uuid":"f10445f9398dcd57b404591c69075a23ab3d5115","_cell_guid":"7e06a571-5ca6-4a50-8bf1-c037d9399df3"}},{"cell_type":"code","source":"fft_all = []\nnames = []\nfor direct in dirs:\n    waves = [f for f in os.listdir(join(train_audio_path, direct)) if f.endswith('.wav')]\n    for wav in waves:\n        sample_rate, samples = wavfile.read(train_audio_path + direct + '/' + wav)\n        if samples.shape[0] != sample_rate:\n            samples = np.append(samples, np.zeros((sample_rate - samples.shape[0], )))\n        x, val = custom_fft(samples, sample_rate)\n        fft_all.append(val)\n        names.append(direct + '/' + wav)\n\nfft_all = np.array(fft_all)\n\n# Normalization\nfft_all = (fft_all - np.mean(fft_all, axis=0)) / np.std(fft_all, axis=0)\n\n# Dim reduction\npca = PCA(n_components=3)\nfft_all = pca.fit_transform(fft_all)\n\ndef interactive_3d_plot(data, names):\n    scatt = go.Scatter3d(x=data[:, 0], y=data[:, 1], z=data[:, 2], mode='markers', text=names)\n    data = go.Data([scatt])\n    layout = go.Layout(title=\"Anomaly detection\")\n    figure = go.Figure(data=data, layout=layout)\n    py.iplot(figure)\n    \ninteractive_3d_plot(fft_all, names)","metadata":{"_uuid":"5a52c1fbea1b68c78502ac9954bd7ecae19149d6","_cell_guid":"8ffe76a4-7e56-481b-8406-22520b573edc","execution":{"iopub.status.busy":"2023-09-09T23:31:40.829287Z","iopub.execute_input":"2023-09-09T23:31:40.829623Z","iopub.status.idle":"2023-09-09T23:33:07.864427Z","shell.execute_reply.started":"2023-09-09T23:31:40.829583Z","shell.execute_reply":"2023-09-09T23:33:07.862218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"yes/e4b02540_nohash_0.wav、go/0487ba9b_nohash_0.wav などのポイントが、残りのポイントから遠く離れたところにあることに注意してください。 聞いてみましょう。\n\nNotice that there are *yes/e4b02540_nohash_0.wav*, *go/0487ba9b_nohash_0.wav* and more points, that lie far away from the rest. Let's listen to them.","metadata":{"_uuid":"f0807b3e37e3ef955abb1995bfa8cccf2d979fd2","_cell_guid":"7c239d25-c761-4c54-87ed-c47c3f641fdd"}},{"cell_type":"code","source":"print('Recording go/00b01445_nohash_0.wav')\n#ファイル無しエラーとなるため別ファイルを指定した\nipd.Audio(join(train_audio_path, 'go/00b01445_nohash_0.wav'))","metadata":{"_uuid":"8f2b077bda2f985ee5dabfa66778d1c3001cb24b","_cell_guid":"70c6eba5-0621-4292-b9bd-b92bc5066ce3","execution":{"iopub.status.busy":"2023-09-09T23:33:07.866880Z","iopub.execute_input":"2023-09-09T23:33:07.867372Z","iopub.status.idle":"2023-09-09T23:33:07.884855Z","shell.execute_reply.started":"2023-09-09T23:33:07.867316Z","shell.execute_reply":"2023-09-09T23:33:07.883990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Recording yes/e4b02540_nohash_0.wav')\nipd.Audio(join(train_audio_path, 'yes/e4b02540_nohash_0.wav'))","metadata":{"_uuid":"4116f932a79f3e2e9f745ec4e441c0529ccf66fa","_cell_guid":"1be3b686-d99d-4a1b-aa23-165ad4e5c67e","execution":{"iopub.status.busy":"2023-09-09T23:33:07.885851Z","iopub.execute_input":"2023-09-09T23:33:07.886242Z","iopub.status.idle":"2023-09-09T23:33:07.901491Z","shell.execute_reply.started":"2023-09-09T23:33:07.886199Z","shell.execute_reply":"2023-09-09T23:33:07.900900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"個々の単語の異常を探す場合は、たとえば次の「7」のファイルを見つけることができます。\n\nIf you will look for anomalies for individual words, you can find for example this file for *seven*:","metadata":{"_uuid":"f1252a107c341e3592a301398be2a56202b9f2a5","_cell_guid":"755472e7-6c19-4c2b-bb8a-d161a501c80a"}},{"cell_type":"code","source":"print('Recording seven/e4b02540_nohash_0.wav')\nipd.Audio(join(train_audio_path, 'seven/b1114e4f_nohash_0.wav'))","metadata":{"_uuid":"542a9b0b0c61e2420af77d0f0196178475952681","_cell_guid":"45552683-e626-4143-88c2-b22f9fa00737","execution":{"iopub.status.busy":"2023-09-09T23:33:07.902400Z","iopub.execute_input":"2023-09-09T23:33:07.902765Z","iopub.status.idle":"2023-09-09T23:33:07.910291Z","shell.execute_reply.started":"2023-09-09T23:33:07.902723Z","shell.execute_reply":"2023-09-09T23:33:07.909752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"それは明らかに重要なことではありません。 通常、この方法を使用すると歪みを見つけることができます。 データには、あるべきものが含まれているようです。\n\nThat's nothing obviously important. Usually you can find some distortions using this method. Data seems to contain what it should.","metadata":{"_uuid":"8e532f1e54b04ba5850a163428d1fd94c47bee37","_cell_guid":"67b4915a-7459-4863-a696-ac8220da5a29"}},{"cell_type":"markdown","source":"## 3. インスピレーションをどこで探すか¶\n\nコンテストではさまざまなアプローチが可能です。 それについては何もアドバイスできません。 私の最初の考えを共有したいと思います。\n\n近年、ニューラルネットワークに基づいたソリューションを提案する傾向にあります。 通常、2 つのアーキテクチャがあります。 私のアイデアはここにあります。\n\nエンコーダ-デコーダ: https://arxiv.org/abs/1508.01211\n\nCTC 損失のある RNN: https://arxiv.org/abs/1412.5567\n私にとって、特に SR 分野の経験がない場合、このコンテストでは 1 と 2 が賢明な選択です。 彼らはエンドツーエンドのソリューションになろうとしています。 音声認識は非常に大きなテーマであり、重要なことを短期間で理解するのは難しいでしょう。\n\n古典的な音声認識については、ここで説明されています: http://www.ece.ucsb.edu/Faculty/Rabiner/ece259/Reprints/tutorial%20on%20hmm%20and%20applications.pdf\n\nカルディのダミー向けチュートリアルがあり、このコンテストと似たような問題がいくつかあります。\n\n非常に深い CNN - SR に使用されるかどうかは不明。 ただし、ほとんどの論文は大語彙連続音声認識システム (LVCSR) に関するものです。 ここでは別のタスクが必要になります。語彙が非常に少なく、単語が 1 つだけ含まれており、(ほとんどの場合) 指定された長さで録音されます。 このようなアプローチが競争に勝つことができると思います。\n\n## 3. Where to look for the inspiration\n<a id=\"wheretostart\"></a> \n\nYou can take many different approches for the competition. I can't really advice any of that. I'd like to share my initial thoughts.\n\nThere is a trend in recent years to propose solutions based on neural networks. Usually there are two architectures. My ideas are here.\n\n1. Encoder-decoder: https://arxiv.org/abs/1508.01211\n2. RNNs with CTC loss: https://arxiv.org/abs/1412.5567<br>\nFor me, 1 and 2  are a sensible choice for this competition, especially if you do not have background in SR field. They try to be end-to-end solutions. Speech recognition is a really big topic and it would be hard to get to know important things in short time.\n\n3. Classic speech recognition is described here: http://www.ece.ucsb.edu/Faculty/Rabiner/ece259/Reprints/tutorial%20on%20hmm%20and%20applications.pdf\n\nYou can find *Kaldi* [Tutorial for dummies](http://kaldi-asr.org/doc/kaldi_for_dummies.html), with a problem similar to this competition in some way.\n\n4. Very deep CNN - Don't know if it is used for SR. However, most papers concern Large Vocabulary Continuous Speech Recognition Systems (LVCSR). We got different task here - a very small vocabulary, and recordings with only one word in it, with a (mostly) given length. I suppose such approach can win the competition. \n","metadata":{"_uuid":"a50ecff1ab0f1629bbd069c5e325e09035bc2778","_cell_guid":"8e1ba6de-b802-417a-b98d-72d0fed93296"}},{"cell_type":"markdown","source":"エンコーダ-デコーダ: https://arxiv.org/abs/1508.01211\n[Submitted on 5 Aug 2015 (v1), last revised 20 Aug 2015 (this version, v2)]\nListen, Attend and Spell\nWilliam Chan, Navdeep Jaitly, Quoc V. Le, Oriol Vinyals\n私たちは、音声発話を文字に書き写すことを学習するニューラル ネットワークである Listen, Attend and Spell (LAS) を紹介します。 従来の DNN-HMM モデルとは異なり、このモデルは音声認識装置のすべてのコンポーネントを共同で学習します。 私たちのシステムには、リスナーとスペラーという 2 つのコンポーネントがあります。 リスナーは、フィルター バンク スペクトルを入力として受け入れるピラミッド型リカレント ネットワーク エンコーダーです。 Speller は、文字を出力として出力する、アテンションベースのリカレント ネットワーク デコーダーです。 ネットワークは、文字間の独立性を仮定せずに文字シーケンスを生成します。 これは、以前のエンドツーエンド CTC モデルに対する LAS の重要な改善点です。 Google 音声検索タスクのサブセットでは、LAS は辞書や言語モデルを使用しない場合に 14.1% の単語誤り率 (WER) を達成し、言語モデルを使用した場合は上位 32 ビームで 10.3% のスコアを達成します。 比較すると、最先端の CLDNN-HMM モデルは 8.0% の WER を達成します。\n\nWe present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent network encoder that accepts filter bank spectra as inputs. The speller is an attention-based recurrent network decoder that emits characters as outputs. The network produces character sequences without making any independence assumptions between the characters. This is the key improvement of LAS over previous end-to-end CTC models. On a subset of the Google voice search task, LAS achieves a word error rate (WER) of 14.1% without a dictionary or a language model, and 10.3% with language model rescoring over the top 32 beams. By comparison, the state-of-the-art CLDNN-HMM model achieves a WER of 8.0%.","metadata":{}},{"cell_type":"markdown","source":"CTC 損失のある RNN: https://arxiv.org/abs/1412.5567\n[Submitted on 17 Dec 2014 (v1), last revised 19 Dec 2014 (this version, v2)]\nエンドツーエンドの深層学習を使用して開発された最先端の音声認識システムを紹介します。 私たちのアーキテクチャは、苦労して設計された処理パイプラインに依存する従来の音声システムよりも大幅にシンプルです。 これらの従来のシステムは、騒音の多い環境で使用するとパフォーマンスが低下する傾向があります。 対照的に、私たちのシステムは、背景雑音、残響、またはスピーカーの変動をモデル化するために手動で設計されたコンポーネントを必要とせず、代わりにそのような影響に対して堅牢な関数を直接学習します。 音素辞書は必要ありませんし、「音素」という概念さえも必要ありません。 私たちのアプローチの鍵となるのは、複数の GPU を使用する適切に最適化された RNN トレーニング システムと、トレーニング用に大量のさまざまなデータを効率的に取得できる一連の新しいデータ合成技術です。 Deep Speech と呼ばれる私たちのシステムは、広く研究されている Switchboard Hub5'00 で以前に公開された結果を上回り、テスト セット全体で 16.0% の誤差を達成しました。 Deep Speech は、広く使用されている最先端の商用音声システムよりも、困難な騒音の多い環境にもうまく対処します。\n\nDeep Speech: Scaling up end-to-end speech recognition\nAwni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, Andrew Y. Ng\nWe present a state-of-the-art speech recognition system developed using end-to-end deep learning. Our architecture is significantly simpler than traditional speech systems, which rely on laboriously engineered processing pipelines; these traditional systems also tend to perform poorly when used in noisy environments. In contrast, our system does not need hand-designed components to model background noise, reverberation, or speaker variation, but instead directly learns a function that is robust to such effects. We do not need a phoneme dictionary, nor even the concept of a \"phoneme.\" Key to our approach is a well-optimized RNN training system that uses multiple GPUs, as well as a set of novel data synthesis techniques that allow us to efficiently obtain a large amount of varied data for training. Our system, called Deep Speech, outperforms previously published results on the widely studied Switchboard Hub5'00, achieving 16.0% error on the full test set. Deep Speech also handles challenging noisy environments better than widely used, state-of-the-art commercial speech systems.","metadata":{}},{"cell_type":"markdown","source":"古典的な音声認識については、ここで説明されています: http://www.ece.ucsb.edu/Faculty/Rabiner/ece259/Reprints/tutorial%20on%20hmm%20and%20applications.pdf","metadata":{}},{"cell_type":"markdown","source":"---\n\n**If you like my work please upvote.**\n\nLeave a feedback that will let me improve! ","metadata":{"_uuid":"f4862ca5ef5f190c0593ff4bd165acc74588c373","_cell_guid":"9192984a-7307-41e5-bb02-9f46f14822c7"}}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}