{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Interactive EDA for BirdClef 2023: Embeddings, baseline results and label errors for windowed data\n![](https://spotlight.renumics.com/resources/birdclef2023.png)\n\n\n# Motivation\n\nFor many use cases, the best way to build robust models is to iterate the training data.\n\nThe BirdClef settings holds one big challenge in the training data: We have labels available for longer recordings, but we are training on shorter windows. One heuristic to deal with this is just use the first 5 seconds of the recording ([discussion](https://www.kaggle.com/competitions/birdclef-2021/discussion/230733)).\n\nWe would like to explore if we can optimize the training data set for the competition. In a first steps, we provide a dataset of windowed samples and several enrichments for these windows (baseline embedding and prediction, label error scores). We hope that this helps to build a great data set, in particular with respect to\n- window selection and sampling\n- oversampling (in particular to combat hidden stratification)\n- discarding low-quality examples\n\n**Feedback and bug reports are highly appreciated**\n\n# Interactive exploration\n\nWe show how the enriched dataset can be explored with our open source data curation tool Spotlight.\n> **If you find the approach useful, we'd be happy if you star the Spotlight repo on Github:** https://github.com/Renumics/spotlight\n\nWe use pyngrok to forward a port from the Kaggle machine to access the Spotlight web application. You need a free [ngrok](https://ngrok.com/) account to use that. Alternatively, you can use similar services such as localtunnel or you can use the approach in your local notebook without port forwarding.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Install required libraries and imports","metadata":{}},{"cell_type":"code","source":"#install required libraries\n!pip install renumics-spotlight pyngrok cleanlab \n\n#imports\nimport numpy as np \nimport pandas as pd \n\nimport getpass\nfrom pyngrok import ngrok, conf\n\nfrom renumics import spotlight\n\nimport tensorflow as tf\nimport tensorflow_hub as hub\nimport tensorflow_io as tfio\nimport csv\nimport io\n\nimport librosa\nfrom tqdm import tqdm\nfrom scipy.stats import entropy\nimport os","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-29T15:30:04.657804Z","iopub.execute_input":"2023-03-29T15:30:04.658501Z","iopub.status.idle":"2023-03-29T15:31:13.511697Z","shell.execute_reply.started":"2023-03-29T15:30:04.658440Z","shell.execute_reply":"2023-03-29T15:31:13.509639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Interactive exploration of pre-computed enriched data\n**Make sure you add the birdclef_2023_enriched dataset to your notebook**\n\n> If you use a local notebook, you can immediatly start Spotlight without port forwarding\n","metadata":{}},{"cell_type":"code","source":"#prepare pyngrok environment\nprint(\"Enter your authtoken, which can be copied from https://dashboard.ngrok.com/auth\")\nconf.get_default().auth_token = getpass.getpass()\n\n# specify a port\nngrok_tunnel = ngrok.connect(5005)\n\n# where we can visit our fastAPI app\nprint('Public URL:', ngrok_tunnel.public_url)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T15:31:13.514671Z","iopub.execute_input":"2023-03-29T15:31:13.515582Z","iopub.status.idle":"2023-03-29T15:31:25.927856Z","shell.execute_reply.started":"2023-03-29T15:31:13.515539Z","shell.execute_reply":"2023-03-29T15:31:25.926008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#load dataframe from file\nDF_ENRICHED_PATH = \"/kaggle/input/birdclef2023-enriched/birdclef2023_enriched.parquet\"\ndf = pd.read_parquet(DF_ENRICHED_PATH)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T15:31:25.929664Z","iopub.execute_input":"2023-03-29T15:31:25.930171Z","iopub.status.idle":"2023-03-29T15:31:26.465637Z","shell.execute_reply.started":"2023-03-29T15:31:25.930129Z","shell.execute_reply":"2023-03-29T15:31:26.464667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Interactive analysis of first window in each recording\n\n> In order to open the Spotlight webapp, you have to click on the **Public URL** that ngrok provides","metadata":{}},{"cell_type":"code","source":"#start interactive analysis\ndf_filtered = df.query(\"first_frame == True\")\nlayout=\"/kaggle/input/birdclef2023-enriched/birdclef2023_layout-small.json\"\nprint('Access the Spotlight instance on this public URL: ', ngrok_tunnel.public_url)\nspotlight.show(df_filtered, port=5005, dtype={'image_url': spotlight.Image, 'filename': spotlight.Audio, 'window': spotlight.Window, 'red_emb':spotlight.Embedding }, layout=layout, wait=True)\n","metadata":{"execution":{"iopub.status.busy":"2023-03-29T15:31:26.467692Z","iopub.execute_input":"2023-03-29T15:31:26.468016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Interactive analysis of windows for critical classes\n> Due to the dataset size this analysis might be slow on older machines. Please allow approx. 3 minutes for the dataset to load (depending on your machine and internet connection).","metadata":{}},{"cell_type":"code","source":"layout=\"/kaggle/input/birdclef2023-enriched/birdclef2023_layout-small.json\"\nspotlight.show(df, port=5005, dtype={'image_url': spotlight.Image, 'filename': spotlight.Audio, 'window': spotlight.Window, 'red_emb':spotlight.Embedding }, layout=layout)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Code to re-compute the enrichment\n> **Make sure you enable GPU acceleration in order to use the code**\n\nWe divide each train and test sample into 5 second windows (no overlap) and compute the following enrichments for each window:\n- Per-window embedding with the baseline model\n- Prediction result, prediction probability and entropy of the softmax-values for the baseline model\n- Label error scores computed with the Cleanlab library\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## Load baseline model and perform inference\nAccording to this [notebook](https://www.kaggle.com/code/philculliton/inferring-birds-with-kaggle-models) by the organizers. ","metadata":{"execution":{"iopub.status.busy":"2023-03-21T14:30:25.439937Z","iopub.execute_input":"2023-03-21T14:30:25.440409Z","iopub.status.idle":"2023-03-21T14:30:25.469763Z","shell.execute_reply.started":"2023-03-21T14:30:25.440349Z","shell.execute_reply":"2023-03-21T14:30:25.468612Z"}}},{"cell_type":"code","source":"model = hub.load('https://kaggle.com/models/google/bird-vocalization-classifier/frameworks/tensorFlow2/variations/bird-vocalization-classifier/versions/1')\nlabels_path = hub.resolve('https://kaggle.com/models/google/bird-vocalization-classifier/frameworks/tensorFlow2/variations/bird-vocalization-classifier/versions/1') + \"/assets/label.csv\"\n\n# Find the name of the class with the top score when mean-aggregated across frames.\ndef class_names_from_csv(class_map_csv_text):\n    \"\"\"Returns list of class names corresponding to score vector.\"\"\"\n    with open(labels_path) as csv_file:\n        csv_reader = csv.reader(csv_file, delimiter=',')\n        class_names = [mid for mid, desc in csv_reader]\n        return class_names[1:]\n\n## note that the bird classifier classifies a much larger set of birds than the\n## competition, so we need to load the model's set of class names or else our \n## indices will be off.\nclasses = class_names_from_csv(labels_path)\nclasses[0]='default'\n\n\n\ntrain_metadata = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\ntrain_metadata.head()\ncompetition_classes = sorted(train_metadata.primary_label.unique())\n\ncompetition_class_map = []\nfor c in competition_classes:\n    try:\n        i = classes.index(c)\n        competition_class_map.append(i)\n    except:\n        competition_class_map.append(0)\n \n\n       \n        ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def ensure_sample_rate(waveform, original_sample_rate,\n                       desired_sample_rate=32000):\n    \"\"\"Resample waveform if required.\"\"\"\n    if original_sample_rate != desired_sample_rate:\n        waveform = tfio.audio.resample(waveform, original_sample_rate, desired_sample_rate)\n    return desired_sample_rate, waveform\n\ndef frame_audio(\n      audio_array: np.ndarray,\n      window_size_s: float = 5.0,\n      hop_size_s: float = 5.0,\n      sample_rate = 32000,\n      ) -> np.ndarray:\n    \n    \"\"\"Helper function for framing audio for inference.\"\"\"\n    \"\"\" using tf.signal \"\"\"\n    if window_size_s is None or window_size_s < 0:\n        return audio_array[np.newaxis, :]\n    frame_length = int(window_size_s * sample_rate)\n    hop_length = int(hop_size_s * sample_rate)\n    framed_audio = tf.signal.frame(audio_array, frame_length, hop_length, pad_end=True)\n    return framed_audio\n\ndef predict_sample(row, window, classes):\n    with tf.device('/gpu:0'):\n        logits, embeddings = model.infer_tf(window[np.newaxis, :])            \n        probabilities = tf.nn.softmax(logits).numpy().flatten()\n    \n    #restrict to classes from competition\n    probabilities=probabilities[competition_class_map]\n    class_id = np.argmax(probabilities)\n    row['prediction']=competition_classes[class_id]\n    row['embedding']=embeddings.numpy().flatten()\n    row['entropy']=entropy(probabilities) \n    row['probabilities']=probabilities","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#load metadata\nDIR_PATH = \"/kaggle/input/birdclef-2023/train_audio/\"\ntrain_df = pd.read_csv('/kaggle/input/birdclef-2023/train_metadata.csv')\ntrain_df['filename'] = train_df['filename'].apply(lambda x: DIR_PATH + x)\n\n#output file\nDF_INFERENCE_PATH = \"/kaggle/working/birdclef_inference.parquet\"\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndf = pd.DataFrame(columns=train_df.columns.values.tolist()+['window', 'energy', 'first_frame', 'prediction','embedding', 'entropy','probabilities'])\n\nglobal_idx = 0\n\n\nfor index,row in tqdm(train_df.iterrows(), total=train_df.shape[0]):\n    \n    audio, sample_rate = librosa.load(row['filename'])\n    sample_rate, wav_data = ensure_sample_rate(audio, sample_rate)  \n    length = librosa.get_duration(y=wav_data, sr = sample_rate)\n\n    \n    windows_idx = np.arange(0, len(wav_data), 5*sample_rate)\n    windows_time = np.arange(0, length, 5)   \n    \n    fixed_tm = frame_audio(wav_data)   \n    \n    current_row = row\n    current_row['first_frame'] = True     \n\n    i=0\n    for i, window in enumerate(fixed_tm[:-1]):\n        #print(str(i))\n        current_row['window']=[windows_time[i], windows_time[i+1]]\n        current_row['energy']=np.sum(wav_data[windows_idx[i]:windows_idx[i+1]]**2) / 5.0\n        \n        predict_sample(row=current_row, window=window, classes=classes)   \n      \n        df.loc[global_idx] = current_row\n        global_idx += 1\n        current_row['first_frame'] = False           \n    \n    #last window    \n    current_row['window']=[windows_time[-1], length]\n    frame_length = len(wav_data)-windows_idx[-1]   \n    current_row['energy']=np.sum(wav_data[windows_idx[-1]:len(wav_data)]**2) / frame_length\n    \n    window = fixed_tm[-1]\n    \n    predict_sample(row=current_row, window=window, classes=classes) \n\n    df.loc[global_idx] = current_row\n    global_idx += 1\n   \n    #save every 1000th sample\n    if (index%1000 == 0):      \n        df.to_parquet(DF_INFERENCE_PATH,compression='None')         \n\n    \ndf.to_parquet(DF_INFERENCE_PATH,compression='None')       \n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Add reduced embeddings\n> It is recommended to re-start the notebook and re-load from file to avoid memory issues during UMAP computation","metadata":{}},{"cell_type":"code","source":"import umap\nimport numpy as np \nimport pandas as pd \nfrom scipy.special import softmax\n\n\n#we load the data \ndf = pd.read_parquet(DF_INFERENCE_PATH)\nDF_RESULT_PATH = \"/kaggle/working/birdclef_enriched.parquet\"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"embeddings = np.stack(df['embedding'].to_numpy())\nreducer = umap.UMAP()\nreduced_embedding = reducer.fit_transform(embeddings)\ndf['red_emb'] = np.array(reduced_embedding).tolist()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#compute count\nlabel_counts = df.value_counts(['primary_label'])\ncount =  [label_counts[x] for x in df['primary_label']]\ndf['count']=count\n\ndf['pred_correct'] = df['prediction'] == df['primary_label']\n\n#add urls\ndf_url = pd.read_json('/kaggle/input/birdclef2023-enriched/image_urls.json')\ndf['image_url']=df_url['image_url']\n\n#drop full embeddings\ndf=df.drop(columns=['embedding'])\n\n#save file\ndf.to_parquet(DF_RESULT_PATH)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Add flags for label errors / difficult examples","metadata":{}},{"cell_type":"code","source":"from cleanlab.filter import find_label_issues \n\n#convert labels into indices\nlabel_index=[competition_classes.index(x) for x in df['primary_label']]\n\n# Find issues and append issue indicators to dataset\n# Also, append column that can be used for relabeling\nlabel_issues = find_label_issues(\n    labels=label_index,\n    pred_probs=np.stack(df.probabilities.values)\n)\n\ndf['label_errors']=label_issues\n\ndf=df.drop(columns=['probabilities'])\n\n#save file\ndf.to_parquet(DF_RESULT_PATH)\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ","metadata":{}}]}