{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"from IPython.display import IFrame","metadata":{"execution":{"iopub.status.busy":"2023-08-17T09:54:35.862176Z","iopub.execute_input":"2023-08-17T09:54:35.862553Z","iopub.status.idle":"2023-08-17T09:54:35.868091Z","shell.execute_reply.started":"2023-08-17T09:54:35.862523Z","shell.execute_reply":"2023-08-17T09:54:35.866483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📊🔍 Exploratory Data Analysis - Bengali.AI Kaggle Competition","metadata":{}},{"cell_type":"markdown","source":"This Notebook aims at doing a brief **exploratory data exploration** of the data available for the Bengali.AI Competition. Concretely by going through this notebook you will get the following:\n1. **Insights** on the available training and validation data (basic information, biases, outliers, duplicates, ...)\n2. A **template for doing interactive EDA** in the EDA tool **[Spotlight](https://github.com/Renumics/spotlight)**\n3. An **embedding- and feature-enriched dataset** as a starting point for your own analysis\n\nBelow you can see an **interactive preview** what you will get when running the notebook. Concretely this is an interactive view on detected **audio issues** in the dataset:","metadata":{}},{"cell_type":"code","source":"IFrame(\"https://renumics-bengaliai-audio-issues.hf.space\", \"100%\", \"700px\")","metadata":{"execution":{"iopub.status.busy":"2023-08-17T09:54:37.501375Z","iopub.execute_input":"2023-08-17T09:54:37.501785Z","iopub.status.idle":"2023-08-17T09:54:37.509505Z","shell.execute_reply.started":"2023-08-17T09:54:37.501752Z","shell.execute_reply":"2023-08-17T09:54:37.508306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**[Check out this Github Repo for more Info on the Tooling used.](https://github.com/Renumics/spotlight)**","metadata":{}},{"cell_type":"markdown","source":"**!!! NOTE**: For running this you have to **ADJUST THE PATH TO YOUR DATASET** in the **Setup section** below and also **RUN THIS LOCALLY**!","metadata":{}},{"cell_type":"markdown","source":"## Setup","metadata":{}},{"cell_type":"code","source":"# Install these dependencies\n!pip install -U sliceguard pandas numpy plotly datasets scikit-learn tqdm renumics-spotlight","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Configure the path to your dataset here\nINPUT_DIR = \"/kaggle/input/bengaliai-speech\" # CHANGE DATASET PATH HERE!!!","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The imports you will need\nfrom pathlib import Path\nfrom tqdm import tqdm\nimport pandas as pd\nimport numpy as np\nimport datasets\nimport plotly.express as px\nfrom sklearn.neighbors import KDTree\nfrom renumics import spotlight\nfrom renumics.spotlight import Audio, Embedding\nfrom sliceguard import SliceGuard","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load the data\nThis code is simply for loading the data. *df* will afterwards contain a feature- and embedding-enriched dataset.","metadata":{}},{"cell_type":"code","source":"# Load the raw data from your machine\ndf = pd.read_csv(Path(INPUT_DIR) / \"train.csv\")\n\n# Pull features and embeddings from huggingface dataset hub\ndataset = datasets.load_dataset(\"renumics/bengaliai-competition-features-embeddings\")\nfeature_df = dataset[\"train\"].to_pandas()\n\n# Merge the two datasets\nadditional_columns = feature_df.columns.difference(df.columns).tolist() + [\"id\"]\n\ndf = pd.merge(df, feature_df[additional_columns], on='id')\n\nif not INPUT_DIR.endswith(\"/\"):\n    INPUT_DIR = INPUT_DIR + \"/\"\ndf[\"audio\"] = INPUT_DIR + df[\"audio\"]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 📈 Basics about the raw data","metadata":{}},{"cell_type":"code","source":"# Sample count and columns\nprint(f\"Sample count is {len(df)}.\")\nprint(f\"Dataframe contains the columns {df.columns.tolist()}.\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dataframe Structure:**\n* The dataframe contains almost 1 million rows\n* It contains the following basic columns:\n    * *id*: The sample id which maps to the corresponding audio file in \"train_mp3s\"\n    * *sentence* The ground-truth transcription of the audio file \n* The enrichment adds the following columns:\n    * *audio_length_s*: Length of the audio file in seconds\n    * *audio_rms_max*: Maximum signal energy of the sample\n    * *audio_rms_mean*: Mean signal energy of the sample\n    * *audio_rms_std*: Maximum signal energy standard deviation\n    * *audio_spectral_flatness_mean*: Audio spectral flatness mean (the higher the more noise like)\n    * *audio_embedding*: Audio embeddings computed using embedding model trained on Audioset\n    * *text_embedding*: Multilingual text embeddings","metadata":{}},{"cell_type":"code","source":"# Split ratio\nprint(\"##### Distribution between splits #####\")\npx.histogram(df, x=\"split\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Split Ratio:**\n* Around 30k samples of all public data are part of the validation set.","metadata":{}},{"cell_type":"code","source":"# Check if the split is random or group-wise\nspotlight.show(df.groupby(\"split\").sample(3000), dtype={\"audio_embedding\": Embedding, \"text_embedding\": Embedding, \"audio\": Audio})","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Which type of split is this?**\n* Seems like a not completely random sample-wise split.\n* Most of the audio embedding space covered by samples from both splits.\n* However, there are groups that only exist in the train split.\n* Additionally there are regions where the validation data is a lot denser, however to some extent exists in the train data.\n* So just beware that it could make sense to track specific regions in your evaluation and potentially adjust the sample distribution if your model is weak with certain types of samples!\n* Something similar is also observable for the text_embedding field, however the embedding quality is probably quite bad so the effect cannot be shown as strongly here.","metadata":{}},{"cell_type":"markdown","source":"## 📊 Distribution of Simple Audio Features","metadata":{}},{"cell_type":"code","source":"print(\"##### Distribution of audio lengths (s) #####\")\naudio_length_fig = px.histogram(df, x=\"audio_length_s\", nbins=200)\naudio_length_fig.show()\n\nprint(\"##### Distribution of rms (signal power) means #####\")\naudio_rms_means_fig = px.histogram(df, x=\"audio_rms_mean\", nbins=200)\naudio_rms_means_fig.show()\n\nprint(\"##### Distribution of rms (signal power) maxs #####\")\naudio_rms_maxs_fig = px.histogram(df, x=\"audio_rms_max\", nbins=200)\naudio_rms_maxs_fig.show()\n\nprint(\"##### Distribution of rms (signal power) stds #####\")\naudio_rms_stds_fig = px.histogram(df, x=\"audio_rms_std\", nbins=200)\naudio_rms_stds_fig.show()\n\n\nprint(\"##### Distribution of spectral flatness (noisyness) #####\")\naudio_spectral_flatness_fig = px.histogram(df, x=\"audio_spectral_flatness_mean\", nbins=200)\naudio_spectral_flatness_fig.show()\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Simple Audio Feature Distributions:**\n* Most of the audio files are around 1.8-6.5 seconds long.\n* There are very few samples below 1.6 seconds and almost none aboce 10.6 seconds length.\n* There seem to exist two peaks in the mean and max distribution of the signal energy, maybe caused by normalizations of two input datasets?\n* The signal energy also has a longer tail towards the high end of the signal energy, meaning there will be a large spectrum of louder samples with varying signal energy.\n* Most of the data seems to be of quite tonal nature. However, there are around 300k samples that are a little noisy and quite a few samples with high noisyness level. (according to spectral flatness)\n\n**For all these findings:** Consider checking if the distribution if the same in train, val or try to experimentally determine if it seems to be the same in test. If not, consider adjusting the distribution, e.g., by normalizing the loudness of the signal.","metadata":{}},{"cell_type":"markdown","source":"## 🔍 Biases and Embedding-based Distributions","metadata":{}},{"cell_type":"markdown","source":"### Interactive Bias Detection","metadata":{}},{"cell_type":"code","source":"spotlight.show(df.sample(10000), dtype={\"audio_embedding\": Embedding, \"text_embedding\": Embedding, \"audio\": Audio})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Embedding-based Bias Detection**:\n* There seems to be a slight bias towards male speakers.\n* ...more biases to be identified via baseline model.","metadata":{}},{"cell_type":"markdown","source":"## 🔄📊 Duplicates and Near Duplicates","metadata":{}},{"cell_type":"code","source":"# Exact duplicates in sentences\nprint(\"##### Distribution on exact duplicates in sentences #####\")\npx.histogram(df[\"sentence\"].value_counts())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Exact duplicates (Text):**\n* There are around 350k sentences that are exactly duplicated in the dataset.\n* 67k samples sentences are duplicated twice.\n* Few sentences are duplicated significantly more.","metadata":{}},{"cell_type":"code","source":"# Near duplicates in sentences\ntext_embeddings = np.vstack(df[\"text_embedding\"].sample(20000))\nkdtree = KDTree(text_embeddings)\n\ndistances = []\nfor emb in tqdm(text_embeddings):\n    dist, ind = kdtree.query([emb], k=10)\n    distances.append(dist[0])\n\ndistances = np.array(distances)\ndistances = distances[:,1:]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.histogram(distances.min(1))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Near Duplicates (Text):**\n* It seems like there are about 2.5% of samples with near or identical duplicates.\n* Then there are very few samples that are maybe really similar, however they are not a significant portion.","metadata":{}},{"cell_type":"code","source":"# Near duplicates in audio\naudio_embeddings = np.vstack(df[\"audio_embedding\"].sample(20000))\nkdtree = KDTree(audio_embeddings)\n\ndistances = []\nfor emb in tqdm(audio_embeddings):\n    dist, ind = kdtree.query([emb], k=10)\n    distances.append(dist[0])\n\ndistances = np.array(distances)\ndistances = distances[:,1:]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.histogram(distances.min(1))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Near Duplicates (Audio):**\n* It seems like there are no distances close to zero.\n* This makes near duplicates being present in the audio data unlikely.","metadata":{}},{"cell_type":"markdown","source":"## 🤨⚠️ Outliers/Anomalies/Errors","metadata":{}},{"cell_type":"code","source":"# Note that for calculating outliers for the full dataset you will need around 40GB of RAM.\n# To decrease the amount of memory needed downsample the data. Of course this will throw away some potential outliers,\n# however enough will be left to get a feel for typical problematic cases.\nNUM_SAMPLES = 30000\n\nif NUM_SAMPLES is not None:\n    selected_indices = np.random.choice(np.arange(len(df)), size=NUM_SAMPLES)\nelse:\n    selected_indices = np.arange(len(df))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Detect larger potential problem clusters via audio embeddings. (min. problem cluster size = 100)\n# The library will essentially fit an outlier detection model and search for clusters of data that are more anomal than their parent clusters.\nsg = SliceGuard()\nissues = sg.find_issues(df.iloc[selected_indices], [\"audio_embedding\"], min_drop=0.03, min_support=100, drop_reference=\"parent\", precomputed_embeddings={\"audio_embedding\": np.vstack(df[\"audio_embedding\"].iloc[selected_indices])})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display only the found issues, no other samples.\n# Remove the parameter if you want to see all data, but beware the interactive report gets jerky above 50k samples.\n_ = sg.report(non_issue_portion=0.0)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Large audio problem clusters (biases/errors)**:\n* There are quite many samples with interrupted audio recordings (probably CPU load or sampling rate problem when recording)\n* There is a large cluster of samples with voice-like background noises which could confuse a model.\n* There are quite many recordings where the speaker speaks especially unclear\n* Some audio recordings are seemingly extremely fast. Maybe saved with a wrong sample rate?\n* There are audio recordings containing a continuous buzzing noise in the background.\n* Some speakers seem to have been recorded with really bad recording equipment.\n* There is some recordings that are not spoken clearly but sung or whispered.","metadata":{}},{"cell_type":"code","source":"# Detect outliers based on the audio_embedding column. (min problem cluster size = 10)\n# The library will essentially fit an outlier detection model and search for clusters of data that are more anomal than their parent clusters.\nsg = SliceGuard()\nissues = sg.find_issues(df.iloc[selected_indices], [\"audio_embedding\"], min_drop=0.02, min_support=3, drop_reference=\"parent\", precomputed_embeddings={\"audio_embedding\": np.vstack(df[\"audio_embedding\"].iloc[selected_indices])})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display only the found issues, no other samples.\n# Remove the parameter if you want to see all data, but beware the interactive report gets jerky above 50k samples.\n_ = sg.report(non_issue_portion=0.0)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Small audio problem clusters (outliers):**\n* There is quite some data that contains loud traffic noise such as honking\n* ...many more weird audio events related to background noise, loudness and so on.","metadata":{}},{"cell_type":"code","source":"# Detect outliers based on the text_embedding column.\n# The library will essentially fit an outlier detection model and search for clusters of data that are more anomal than their parent clusters.\nsg = SliceGuard()\nsg.find_issues(df, [\"text_embedding\"], min_drop=0.1, min_support=10, precomputed_embeddings={\"text_embedding\": np.vstack(df[\"text_embedding\"])}, drop_reference=\"parent\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display only the found issues, no other samples.\n# Remove the parameter if you want to see all data, but beware the interactive report gets jerky above 50k samples.\nsg.report(non_issue_portion=0.0)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Outliers/Errors/Anomalies (Text):**\n* I do not speak Bengali so I cannot really interpret what I see, so let me know if you do ;-)","metadata":{}},{"cell_type":"markdown","source":"# 🕵️‍♂️ Free EDA in Spotlight\nNow it's your turn. Use Spotlight to uncover even more hidden patterns.","metadata":{}},{"cell_type":"code","source":"spotlight.show(df, dtype={\"audio_embedding\": Embedding, \"text_embedding\": Embedding, \"audio\": Audio})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For more info on how to use the tool **[CHECK OUT THE GITHUB REPOSITORY](https://github.com/Renumics/spotlight)**","metadata":{}}]}