{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Interactive Visualization, Features, Embeddings","metadata":{}},{"cell_type":"markdown","source":"![Spotlight](https://github.com/Renumics/sliceguard/blob/main/static/img/bengaliai_spotlight.png?raw=true)","metadata":{}},{"cell_type":"markdown","source":"This notebook provides you with two resources that can help you to suceed in this competition:\n1. An **enriched dataset version** containing **audio features, as well as audio- and text embeddings**\n2. Code for **interactive exploration** in the data curation tool [Spotlight](https://github.com/Renumics/spotlight) to conduct your own **EDA and Evaluation**\n\nNote that in order for the interactive exploration to work you should **RUN THIS LOCALLY**, not in the kaggle environment.\n\nThe dataset contains the following columns:\n* *audio_length_s*: Length of the audio file in seconds\n* *audio_rms_max*: Maximum signal energy of the sample\n* *audio_rms_mean*: Mean signal energy of the sample\n* *audio_rms_std*: Maximum signal energy standard deviation\n* *audio_spectral_flatness_mean*: Audio spectral flatness mean\n* *audio_embedding*: Audio embeddings computed using embedding model trained on Audioset\n* *text_embedding*: Multilingual text embeddings","metadata":{}},{"cell_type":"code","source":"# IMPORTANT (!): Change this if you are executing this locally for interactive exploration.\n# Set your directory containing the train.csv file\nINPUT_DIR = \"/kaggle/input/bengaliai-speech\"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Install these dependencies to get and view the enriched dataset\n!pip install -U pandas datasets renumics-spotlight==1.3.0rc6","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# All imports\nfrom pathlib import Path\nimport pandas as pd\nimport datasets\nfrom renumics import spotlight\nfrom renumics.spotlight import Audio, Embedding","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the raw data from your machine\ndf = pd.read_csv(Path(INPUT_DIR) / \"train.csv\")\n\n# Pull features and embeddings from huggingface dataset hub\ndataset = datasets.load_dataset(\"renumics/bengaliai-competition-features-embeddings\")\nfeature_df = dataset[\"train\"].to_pandas()\n\n# Merge the two datasets\nadditional_columns = feature_df.columns.difference(df.columns).tolist() + [\"id\"]\ndf = pd.merge(df, feature_df[additional_columns], on='id')\nif not INPUT_DIR.endswith(\"/\"):\n    INPUT_DIR = INPUT_DIR + \"/\"\ndf[\"audio\"] = INPUT_DIR + df[\"audio\"]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the dataframe containing additional features and embeddings\ndf","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Open the dataset for interactive exploration (Subsampled to 5000 samples)\nspotlight.show(df.sample(5000), dtype={\"audio\": Audio, \"audio_embedding\": Embedding, \"text_embedding\": Embedding})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**[Docs for the Exploration Tool on Github](https://github.com/Renumics/spotlight)**","metadata":{}}]}