{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Intro\n\n\n**This notebook's primary purpose is to get an _intuition_ about what factors affect model accuracy.**\n\nThe default model already gives us a very good baseline. Our best bet to improve on it is to focus on the incorrect predictions. For statistical power, we can use the vast training dataset. So we'll run the model on the training data and see what factors / parameters are linked to incorrect predictions.\n\n<br>\n\nBecause of Kaggle limitations, I could only compute model predictions locally. I have uploaded results in a csv here. \n\n_Disclaimer: I did not look through published notebooks before submission, so some ideas may duplicate those of other participants._\n\nAlso this is my first ever submitted notebook, so I am not sure it is going to work as intended. Let me know in the comments if it doesn't.\n\n<br>\n\n\n\n**This notebook has two parts:**<a class='anchor' id='top'></a>\n\n1. [<b>Basic EDA on metadata.</b>](#basic_eda)\n\n    * [Bird classes (labels)](#labels): primary, secondary.\n    \n    * __[Consistency](#consistency) of frequencies__ of primary/secondary labels.\n    \n    * [Ratings](#ratings).\n    \n    * Ratings are [lowered](#lowered) by secondary labels.\n\n2. [<b>Model applied to training data.</b>](#model)\n\n    * [Load model predictions or run the model locally/on kaggle](#apply)\n    \n    * Have the following ready: (1) most probable label, (2) top 10 most probable labels, (3) probabilities, (4) ratings. \n    \n    * [Basic training set stats](#basic_model): model accuracy, ratings for correct/incorrect predictions, probabilities vs correct/incorrect.\n    \n    * [__Correlation of model certainty (logits) and rating.__](#probability_vs_rating)\n    \n    * [Confusion matrix](#confmat).\n    \n    * [__True positive rate, False positive rate.__](#tpr_fpr)\n    \n    * [__Labels with highest accuracy, labels with lowest accuracy.__](#train_accuracy)\n    \n    * [Model finds secondary labels in 1251 of 2305 cases](#top10), based on top 10 predicted labels.\n\n","metadata":{}},{"cell_type":"markdown","source":"Edits:\n\n\n04.04.23 - formatting and legibility.","metadata":{}},{"cell_type":"markdown","source":"# 1. Basic EDA on metadata. <a class='anchor' id='basic_eda'></a> [↑](#top)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ntrain_metadata = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\ntrain_metadata.head(15)","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:32.698087Z","iopub.execute_input":"2023-04-04T03:52:32.699091Z","iopub.status.idle":"2023-04-04T03:52:32.820902Z","shell.execute_reply.started":"2023-04-04T03:52:32.699017Z","shell.execute_reply":"2023-04-04T03:52:32.819422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There's a bunch of potentially useful information. Besides the main bird on the record (<code>'primary_label'</code>) there seem to be other birds (<code>'secondary_labels'</code>), sometimes more than one. There's also a few types of vocalizations (<code>'song'</code>,<code>'call'</code>,etc) which means there can be completely different vocalizations for the same species. Also interestingly, all recordings are given a rating, I'm assuming it measures the quality of the main bird's vocalization, or audio quality, or both. \n\nLet's get some basic stats.","metadata":{}},{"cell_type":"markdown","source":"#### Bird classes (labels)\n\n* Total unique primary labels.\n* Total unique secondary labels.\n* Are all secondary labels among primary, or are there unknown labels?\n* Number of samples per label.\n* Is the bird rarity/occurrence similar in primary and secondary observations?","metadata":{}},{"cell_type":"markdown","source":"#### Ratings\n\n* Distribution of ratings across samples.\n* Does the presence of secondary labels affect the rating?\n* Distribution of average rating across primary labels.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set_context('poster')\nsns.set()","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:32.822538Z","iopub.execute_input":"2023-04-04T03:52:32.823710Z","iopub.status.idle":"2023-04-04T03:52:33.319953Z","shell.execute_reply.started":"2023-04-04T03:52:32.823669Z","shell.execute_reply":"2023-04-04T03:52:33.318898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Bird classes (labels)<a class='anchor' id='labels'></a> [↑](#top)\n\nLet's see how many unique primary and secondary labels we have, and how many secondary labels are among primary.","metadata":{}},{"cell_type":"code","source":"# primary classes:\ncompetition_classes = train_metadata.primary_label.unique()\n\n# accumulate all secondary classes, keep unique:\nall_secondary = []\nfor x in train_metadata.secondary_labels.unique():\n    all_secondary += eval(x)\nall_secondary_unq = list(set(all_secondary))\n\n# number of secondary among primary:\nprint('unique secondary birds detected, total n={0}'.format(len(all_secondary_unq))) \nprint('unique secondary birds in primary (competition classes), total n={0}'.format(len( set(all_secondary_unq) & set(competition_classes) ))) \nprint('total primary classes n={0}'.format(len(competition_classes)))","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:33.321823Z","iopub.execute_input":"2023-04-04T03:52:33.322503Z","iopub.status.idle":"2023-04-04T03:52:33.339864Z","shell.execute_reply.started":"2023-04-04T03:52:33.322462Z","shell.execute_reply":"2023-04-04T03:52:33.338448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of sample counts across primary and secondary labels\n\n# number of samples per primary class:\nplt.figure(figsize=(20,5))\ntrain_metadata.primary_label.value_counts().sort_values(ascending=False).iloc[:100].plot(kind='bar')\nplt.title(' number of samples per primary class (top 100)' )\n\n# number of samples per secondary class:\nplt.figure(figsize=(20,5))\nsecondary_series=pd.Series(all_secondary)\nsecondary_series.value_counts().iloc[:100].sort_values(ascending=False).plot(kind='bar')\nplt.title(' number of samples per secondary class (top 100)' )","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:33.344289Z","iopub.execute_input":"2023-04-04T03:52:33.344715Z","iopub.status.idle":"2023-04-04T03:52:40.095679Z","shell.execute_reply.started":"2023-04-04T03:52:33.344676Z","shell.execute_reply":"2023-04-04T03:52:40.094398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Label consistency<a class='anchor' id='consistency'></a> [↑](#top)\n\nBirds that frequently occur as primary labels, should also frequently be among secondary labels. Rare birds will be rare among both primary and secondary labels. \n\n\nDoes this hold in our data?","metadata":{}},{"cell_type":"code","source":"# Is the commonality of primary and secondary birds consistent?\n\nsecondary_value_counts = secondary_series.value_counts()\nsecondary_value_counts.name = 'secondary'\nprimary_value_counts = train_metadata.primary_label.value_counts()\nprimary_value_counts.name = 'primary'\n\nout = pd.merge(secondary_value_counts, primary_value_counts, left_index=True, right_index=True)\n\nimport numpy as np\nout['log_sec'] = np.log(out['secondary'])\nout['log_prim'] = np.log(out['primary'])\n\nplt.figure()\nsns.scatterplot(data=out, x='secondary', y='primary')\nplt.title('number of occurrences as primary vs secondary classes')\n\nplt.figure()\nsns.scatterplot(data=out, x='log_sec', y='log_prim')\nplt.title('number of occurrences as primary vs secondary classes, log scale \\n red line is a linear regression line')\n\nfrom sklearn.linear_model import LinearRegression\nlr = LinearRegression()\nlr.fit(out['log_sec'][:,np.newaxis], out['log_prim'][:,np.newaxis])\nxregr = np.array([0,5])[:,np.newaxis]\nypred = lr.predict(xregr)\nplt.plot(xregr,ypred,'r')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:40.096893Z","iopub.execute_input":"2023-04-04T03:52:40.097242Z","iopub.status.idle":"2023-04-04T03:52:41.064239Z","shell.execute_reply.started":"2023-04-04T03:52:40.097208Z","shell.execute_reply":"2023-04-04T03:52:41.062900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary:** As expected, birds that are frequently detected as primary class are also frequently detected as secondary. \n\nThis is except a handful of birds with very frequent primary records and very few secondary. ~~Perhaps their vocalizations are just too loud to be secondary?~~ Could the data for these species be collected in a more intentional/targeted way where recording conditions and the environment are different from the rest of the dataset (note these birds have 500 samples each). If this is the case, the audio dataset is heterogeneous, which may be important for the model.","metadata":{}},{"cell_type":"markdown","source":"#### Ratings<a class='anchor' id='ratings'></a> [↑](#top)\n\nEvery recording has a rating, which is assigned by an expert labeller. \n\nFirst let's look at:\n\n- The distribution of ratings across recordings.\n\n- Average ratings for every species. \n\nNext, let's see what could influence a rating. ","metadata":{}},{"cell_type":"code","source":"plt.figure()\ntrain_metadata.rating.value_counts().sort_index().plot(kind='bar')\nplt.title('rating per record')\nplt.xlabel('rating')\nplt.ylabel('n records')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:41.065625Z","iopub.execute_input":"2023-04-04T03:52:41.066099Z","iopub.status.idle":"2023-04-04T03:52:41.382361Z","shell.execute_reply.started":"2023-04-04T03:52:41.066051Z","shell.execute_reply":"2023-04-04T03:52:41.381020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20,5))\ntrain_metadata.groupby('primary_label').rating.mean().plot(kind='bar')\nplt.title('average rating for every primary species')\nplt.ylabel('average rating')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:41.384100Z","iopub.execute_input":"2023-04-04T03:52:41.384460Z","iopub.status.idle":"2023-04-04T03:52:49.944843Z","shell.execute_reply.started":"2023-04-04T03:52:41.384426Z","shell.execute_reply":"2023-04-04T03:52:49.943464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary**: the majority of average ratings are within the interval 3.0-4.0, but there are outliers: birds with unusually high or low ratings. Could the model be improved by looking at these low-ranking outiers in more detail?","metadata":{}},{"cell_type":"markdown","source":"#### Ratings are lowered by secondary labels. <a class='anchor' id='lowered'></a> [↑](#top)\nSo what can influence a rating? \n\nThere can be many different factors: audio quality, type of vocalization (song/call), etc.\n\nOne simple factor to check is the presence of secondary vocalizations. A recording of the primary species may be masked by vocalizations of secondary species. So it is possible that for non-empty secondary labels the rating is lower than on average.","metadata":{}},{"cell_type":"code","source":"# Let's create additional columns for the analysis; we are going to compare the rating of each recording to \n# the average rating of the same primary species only, not the total average rating\n\ntrain_metadata['avg_rating'] = train_metadata.groupby('primary_label').rating.transform('mean')\ntrain_metadata['rating_diff'] = train_metadata['rating'] - train_metadata['avg_rating']\ntrain_metadata['has_secondary'] = train_metadata['secondary_labels']!='[]'\n\ntrain_metadata.head(20)","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:49.946252Z","iopub.execute_input":"2023-04-04T03:52:49.946615Z","iopub.status.idle":"2023-04-04T03:52:49.987915Z","shell.execute_reply.started":"2023-04-04T03:52:49.946554Z","shell.execute_reply":"2023-04-04T03:52:49.986461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# is there a difference between clean recordings and recordings with secondary species?\n\nprint( '\\nrating in records WITH secondary labels differs from average rating by {0} with standard deviation {1}'.format(\n        train_metadata[train_metadata['has_secondary']==True].rating_diff.mean(),\n        train_metadata[train_metadata['has_secondary']==True].rating_diff.std()) \n     )\nprint('\\n')      \nprint( 'rating in records WITHOUT secondary labels differs from average rating by {0} with standard deviation {1}'.format(\n        train_metadata[train_metadata['has_secondary']==False].rating_diff.mean(), \n        train_metadata[train_metadata['has_secondary']==False].rating_diff.std())\n     )","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:49.989328Z","iopub.execute_input":"2023-04-04T03:52:49.989701Z","iopub.status.idle":"2023-04-04T03:52:50.011538Z","shell.execute_reply.started":"2023-04-04T03:52:49.989664Z","shell.execute_reply":"2023-04-04T03:52:50.009808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary:** There is a difference, similar to what we predicted earlier: secondary birds mask the primary one, and reduce the rating of a recording. \n\nHowever, the standard deviation is really high, and it seems that the difference in mean is going to be insignificant if tested by t-test. The difference might turn out significant with a Wilcoxon signed rank test for matched samples (did not explore here).","metadata":{}},{"cell_type":"markdown","source":"# 2. Model applied to training data. <a class='anchor' id='model'></a> [↑](#top)\n\nNow let's run the default model on the training dataset!\n\nI couldn't do it in a Kaggle notebook, so instead I did it locally.\n\nFor every audio file, I got the model logits (probabilities) for every 5s chunk. \n\nI set the label of the whole audio file to the one with the highest probability across all 5s chunks.\n\nThis was my predicted primary label. I also kept the top 10 logits for further analysis, e.g. related to secondary labels. \n\nIn summary, __for this analysis I created a dataframe that has <u>model prediction</u> (most probable label), and <u>top 10 most probable labels</u>__.\n\nWith this data available, there are a few questions that might help us better understand model performance. \n\n<br>\n\n* Does model performance improve with rating?\n\n* How well does the model perform on primary labels? on secondary labels?\n\n* What labels are most frequently confused (confusion matrix).\n\n* False positives and negatives.","metadata":{}},{"cell_type":"markdown","source":"#### Make predictions on the training set<a class='anchor' id='apply'></a> [↑](#top)\n\n\nHere you can either load the dataframe 'model_train_set.csv', or try and calculate this dataframe yourself by running the model locally (or on kaggle, but that takes too long in my experience). \n\n- To load the csv and proceed with the analysis (recommended), run the [Prerequisites](#prereq), and then [Load the dataframe](#load_df).\n\n- To run the code yourself locally, set <code>train_dir = LOCAL_TRAIN_DIR</code>, run the [Prerequisites](#prereq), then [uncomment and run the next cell](#runs).\n\n- To run the code yourself on kaggle, set <code>train_dir = '/kaggle/input/birdclef-2023/train_audio/'</code>, run the [Prerequisites](#prereq), then [uncomment and run the next cell](#runs).","metadata":{}},{"cell_type":"markdown","source":"#### Prerequisites copied from the startup notebook: <a class='anchor' id='prereq'></a> [↑](#apply)\n\nThis basic codeis copied from the startup notebook. It makes a prediction for every 5s sample. ","metadata":{}},{"cell_type":"code","source":"import tensorflow_io as tfio\nimport numpy as np\nimport pandas as pd\nimport tensorflow_hub as hub\nimport tensorflow as tf\nimport librosa\nimport os\nimport csv\n\nmodel = hub.load('https://kaggle.com/models/google/bird-vocalization-classifier/frameworks/tensorFlow2/variations/bird-vocalization-classifier/versions/1')\n\ndef frame_audio(\n      audio_array: np.ndarray,\n      window_size_s: float = 5.0,\n      hop_size_s: float = 5.0,\n      sample_rate = 32000,\n      ) -> np.ndarray:\n    \n    \"\"\"Helper function for framing audio for inference.\"\"\"\n    \"\"\" using tf.signal \"\"\"\n    if window_size_s is None or window_size_s < 0:\n        return audio_array[np.newaxis, :]\n    frame_length = int(window_size_s * sample_rate)\n    hop_length = int(hop_size_s * sample_rate)\n    framed_audio = tf.signal.frame(audio_array, frame_length, hop_length, pad_end=True)\n    return framed_audio\n\n\ndef predict_for_sample(filename, sample_submission, frame_limit_secs=None):\n    file_id = filename.split(\".ogg\")[0].split(\"/\")[-1]\n    print(file_id)\n    \n    audio, sample_rate = librosa.load(filename)\n    sample_rate, wav_data = ensure_sample_rate(audio, sample_rate)\n    \n    fixed_tm = frame_audio(wav_data)\n    \n    frame = 5\n    all_logits, all_embeddings = model.infer_tf(fixed_tm[:1])\n    for window in fixed_tm[1:]:\n        if frame_limit_secs and frame > frame_limit_secs:\n            continue\n        \n        logits, embeddings = model.infer_tf(window[np.newaxis, :])\n        all_logits = np.concatenate([all_logits, logits], axis=0)\n        frame += 5\n    \n    frame = 5\n    all_probabilities = []\n    for frame_logits in all_logits:\n        probabilities = tf.nn.softmax(frame_logits).numpy()\n        \n        ## change sample submission by adding new rows\n        this_index = len(sample_submission.index)\n        print(this_index)\n        sample_submission.loc[this_index,'row_id'] = file_id + \"_\" + str(frame)\n        sample_submission.loc[this_index,competition_classes] = probabilities[competition_class_map]\n        frame += 5        \n\n\ndef ensure_sample_rate(waveform, original_sample_rate,\n                       desired_sample_rate=32000):\n    \"\"\"Resample waveform if required.\"\"\"\n    if original_sample_rate != desired_sample_rate:\n        waveform = tfio.audio.resample(waveform, original_sample_rate, desired_sample_rate)\n    return desired_sample_rate, waveform        \n\n\n\n\n\n\nlabels_path = hub.resolve('https://kaggle.com/models/google/bird-vocalization-classifier/frameworks/tensorFlow2/variations/bird-vocalization-classifier/versions/1') + \"/assets/label.csv\"\n\n# Find the name of the class with the top score when mean-aggregated across frames.\ndef class_names_from_csv(class_map_csv_text):\n    \"\"\"Returns list of class names corresponding to score vector.\"\"\"\n    with open(labels_path) as csv_file:\n        csv_reader = csv.reader(csv_file, delimiter=',')\n        class_names = [mid for mid, desc in csv_reader]\n        return class_names[1:]\n\n## note that the bird classifier classifies a much larger set of birds than the\n## competition, so we need to load the model's set of class names or else our \n## indices will be off.\nclasses= class_names_from_csv(labels_path)\n\n\n\n\n\ntrain_metadata = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\n\n# get competition classes, and class mapping\ncompetition_classes = sorted(train_metadata.primary_label.unique())\nforced_defaults = 0\ncompetition_class_map = []\nfor c in competition_classes:\n    try:\n        i = classes.index(c)\n        competition_class_map.append(i)\n    except:\n        competition_class_map.append(0)\n        forced_defaults += 1\n\n\n\nsample_sub = pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\nsample_sub[competition_classes] = sample_sub[competition_classes].astype(np.float32)\nsample_sub = sample_sub.drop(index=[0,1,2]).reset_index()\nsample_sub = sample_sub.drop('index', axis=1)\n\ndf_model_train = pd.DataFrame(columns=['file','primary','prediction','is_correct','prob','top10'])","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:52:50.013491Z","iopub.execute_input":"2023-04-04T03:52:50.013920Z","iopub.status.idle":"2023-04-04T03:53:00.562994Z","shell.execute_reply.started":"2023-04-04T03:52:50.013882Z","shell.execute_reply":"2023-04-04T03:53:00.561597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Code that runs the model on the training data: <a class='anchor' id='runs'></a> [↑](#apply)\n\nUncomment and run if you want to run the model yourself.","metadata":{}},{"cell_type":"code","source":"\"\"\"for idir,directory in enumerate(os.listdir(train_dir)):\n    print('directory {0} of {1}'.format(idir, len(os.listdir(train_dir))))\n    for file in os.listdir( train_dir+directory+'/' ):\n        filepath = train_dir + directory + '/' + file\n        \n        sample_sub_new = sample_sub.copy()\n        predict_for_sample(filepath, sample_sub_new)\n        \n        top10 = sample_sub_new.loc[:,'abethr1':].stack().sort_values(ascending=False)[:10]\n        \n        df_model_train.loc[len(df_model_train)] = [file,\n                                                                       directory,\n                                                                       top10.index[0][1],\n                                                                       directory == top10.index[0][1],\n                                                                       top10.max(),\n                                                                       top10\n                                                                      ]\n        \n        print('max probability class {0} with probability {1}'.format(top10.index[0][1], top10.max()) )\n        print('file {0} maxprob class matches the original folder: {1}'.format( file, directory == top10.index[0][1]) )\"\"\"\n\n# save to csv\n#df_model_train.to_csv('model_train_set.csv')\n\n# postprocessing (same as in \"Load the dataframe\" cell)\n#df_model_train['top10'] = df_model_train['top10'].apply(lambda x: x.split('\\n'))\n#df_model_train['top10'] = df_model_train['top10'].apply(lambda y: [(x.split()[-2],x.split()[-1]) for x in y[:-1]])\n#df_model_train['top10_labels'] = df_model_train['top10'].apply(lambda y: [x[0] for x in y])\n#df_model_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:00.564701Z","iopub.execute_input":"2023-04-04T03:53:00.565097Z","iopub.status.idle":"2023-04-04T03:53:00.575685Z","shell.execute_reply.started":"2023-04-04T03:53:00.565045Z","shell.execute_reply":"2023-04-04T03:53:00.574281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Load the dataframe  <a class='anchor' id='load_df'></a> [↑](#apply)\n\nLoad the model outputs that I have computed locally. \n\nThe dataset \"bird-v1-train-set\" should be attached to the notebook, it contains model_train_set.csv where model predictions are stored.\n\nApologies for the three lines of code to postprocess model output into convenient columns.","metadata":{}},{"cell_type":"code","source":"df_model_train = pd.read_csv('/kaggle/input/bird-v1-train-set/model_train_set.csv').drop(['Unnamed: 0'],axis=1)\n\ndf_model_train['top10'] = df_model_train['top10'].apply(lambda x: x.split('\\n'))\ndf_model_train['top10'] = df_model_train['top10'].apply(lambda y: [(x.split()[-2],x.split()[-1]) for x in y[:-1]])\ndf_model_train['top10_labels'] = df_model_train['top10'].apply(lambda y: [x[0] for x in y])\n\n\ndf_model_train.head()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2023-04-04T03:53:00.579429Z","iopub.execute_input":"2023-04-04T03:53:00.580922Z","iopub.status.idle":"2023-04-04T03:53:00.898466Z","shell.execute_reply.started":"2023-04-04T03:53:00.580873Z","shell.execute_reply":"2023-04-04T03:53:00.897241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Basic stats: model performance vs ratings and probabilities. <a class='anchor' id='basic_model'></a> [↑](#top)\n\n- How does the model do overall?\n\n- What is the average rating of correct predictions? Incorrect predictions?\n\n- What is the average probability of correct predictions? Incorrect predictions?","metadata":{}},{"cell_type":"code","source":"df_model_train['filename'] = df_model_train['primary']+'/'+df_model_train['file']\ndf_model_train = pd.merge(df_model_train, train_metadata[['filename','rating']], on=['filename'])\n\ndf_model_train['accuracy'] = df_model_train.groupby('primary').is_correct.transform('mean')\n\n\nprint('model accuracy {0}\\n'\n      'average rating of correct detections {1}\\n'\n      'average rating of incorrect detections {2}\\n'\n      'mean probability of correct detections {3}\\n'\n      'mean probability of incorrect detections {4}\\n'\n      .format(df_model_train['is_correct'].mean(),\n              df_model_train[df_model_train['is_correct']].rating.mean(), \n              df_model_train[~df_model_train['is_correct']].rating.mean(),\n              df_model_train[df_model_train['is_correct']].prob.mean(),\n              df_model_train[~df_model_train['is_correct']].prob.mean()\n             )\n     )","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:00.902451Z","iopub.execute_input":"2023-04-04T03:53:00.902827Z","iopub.status.idle":"2023-04-04T03:53:00.951104Z","shell.execute_reply.started":"2023-04-04T03:53:00.902793Z","shell.execute_reply":"2023-04-04T03:53:00.949811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary:** Model works better on records with a higher rating.","metadata":{}},{"cell_type":"markdown","source":"#### Model certainty (logits) vs rating  <a class='anchor' id='probability_vs_rating'></a> [↑](#top)\n\n\n\nProbabilities given by the model could be closely related to the rating.\n\nFor example, bad audio quality could negatively affect both the human rating and the model's certainty.\n\nHowever from a simple scatter plot, this is not obvious:","metadata":{}},{"cell_type":"code","source":"plt.scatter(df_model_train['prob'],df_model_train['rating'])\n\nplt.title('Rating as a function of model probability (logits)')\nplt.xlabel('probability (logits)')\nplt.ylabel('rating')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:00.952601Z","iopub.execute_input":"2023-04-04T03:53:00.952949Z","iopub.status.idle":"2023-04-04T03:53:01.344132Z","shell.execute_reply.started":"2023-04-04T03:53:00.952916Z","shell.execute_reply":"2023-04-04T03:53:01.342655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### KDE version of the same plot\n\nKDE plot below is more informative, and there is a <u>slight positive tilt of the orange cloud</u>, so <u>the correlation between rating and probability may be present</u>. \n    \nA few more interesting observations:\n\n* Model probabilities mostly cluster around 0 and 1, and less around the intermediate values.\n\n* There's a small subset of correctly detected species with rating 0, for which the model was very certain.","metadata":{}},{"cell_type":"code","source":"df_model_train['rating_plot'] = df_model_train.rating + 0.3*np.random.randn(len(df_model_train))\n\nsns.kdeplot(\n    data=df_model_train, x=\"prob\", y=\"rating_plot\", hue=\"is_correct\", fill=False,\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:01.345951Z","iopub.execute_input":"2023-04-04T03:53:01.346410Z","iopub.status.idle":"2023-04-04T03:53:11.486632Z","shell.execute_reply.started":"2023-04-04T03:53:01.346371Z","shell.execute_reply":"2023-04-04T03:53:11.485147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary:** There is a weak correlation between model confidence and rating.","metadata":{}},{"cell_type":"code","source":"# fyi, the rating histogram at 0 has a peak\n\ntrain_metadata['rating'].plot(kind='hist')\nplt.xlabel('rating')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:11.488465Z","iopub.execute_input":"2023-04-04T03:53:11.488918Z","iopub.status.idle":"2023-04-04T03:53:11.734150Z","shell.execute_reply.started":"2023-04-04T03:53:11.488877Z","shell.execute_reply":"2023-04-04T03:53:11.732868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Confusion matrix <a class='anchor' id='confmat'></a> [↑](#top)\n\nSeeing what categories get confused most frequently is another valuable piece of information to improve the model. ","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\n\ncf_labels = df_model_train.primary.unique()\ncf_matrix = confusion_matrix(df_model_train['primary'], df_model_train['prediction'], labels=cf_labels)\n\nsns.heatmap(np.log(np.log(np.log(cf_matrix+1)+1)+1))\nplt.title('confusion matrix')\nplt.xlabel('identified as label')\nplt.ylabel('true label')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:11.735790Z","iopub.execute_input":"2023-04-04T03:53:11.736226Z","iopub.status.idle":"2023-04-04T03:53:12.526780Z","shell.execute_reply.started":"2023-04-04T03:53:11.736189Z","shell.execute_reply":"2023-04-04T03:53:12.525504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary:** Some labels seem to be often falsely detected, some are often incorrectly labeled. This should be reflected in false-positive rates and true positive rates.","metadata":{}},{"cell_type":"markdown","source":"#### False positives, true positives. <a class='anchor' id='tpr_fpr'></a> [↑](#top)\n\nWe can unpack the confusion matrix by checking if there is a category X that is often classified as something else (false negative).\n\nOr if there is a category Y that is incorrectly detected in various examples that are not in fact Y (false positive). \n\nWe can summarize these things in false positive rate and true positive rate for every category.","metadata":{}},{"cell_type":"code","source":"# hit rate aka true positive rate aka recall\n\ntpr = np.zeros((len(cf_matrix),))\nfpr = np.zeros((len(cf_matrix),))\n\n\nfor i in range(len(cf_matrix)):\n    \n    tpr[i] = cf_matrix[i,i]  / np.sum(cf_matrix[i,:])\n    fpr[i] = ( np.sum(cf_matrix[:,i])-cf_matrix[i,i] ) / np.sum(cf_matrix[:,i])\n    \nplt.plot(tpr)\nplt.title('true positive rate (blue); false positive rate (orange)')\nplt.xlabel('label')\nplt.ylabel('proportion')\n\nplt.plot(fpr)\nplt.xlabel('label')\nplt.ylabel('proportion')","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:12.528369Z","iopub.execute_input":"2023-04-04T03:53:12.529021Z","iopub.status.idle":"2023-04-04T03:53:12.806882Z","shell.execute_reply.started":"2023-04-04T03:53:12.528966Z","shell.execute_reply":"2023-04-04T03:53:12.805489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary:** TPR/FPR plot shows a few outliers: birds with basically no correct detections (TPR=0), or high false positive rates. These can be used as pointers for model improvement.","metadata":{}},{"cell_type":"markdown","source":"#### Labels for which the model is most/least accurate.  <a class='anchor' id='train_accuracy'></a> [↑](#top)\n\nTo be more specific, let's look at the labels where the model is the least / most accurate.","metadata":{}},{"cell_type":"code","source":"print('20 most accurate predictions\\n')\nprint( df_model_train.groupby('primary').accuracy.median().sort_values(ascending=False)[:20] )\nprint('\\n\\n')\nprint('20 least accurate predictions\\n')\nprint( df_model_train.groupby('primary').accuracy.median().sort_values(ascending=True)[:20] )","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:12.808236Z","iopub.execute_input":"2023-04-04T03:53:12.808598Z","iopub.status.idle":"2023-04-04T03:53:12.825237Z","shell.execute_reply.started":"2023-04-04T03:53:12.808545Z","shell.execute_reply":"2023-04-04T03:53:12.824047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Secondary labels are found in top 10!  <a class='anchor' id='top10'></a> [↑](#top)\n\nA human expert can sometimes detect secondary vocalizations. Can the model do the same?\n\nWe are going to look at labels in 'top10' detected vocalizations and see if they contain any of the secondary labels.\n\nLet's say that for the model to be successful we need at least one of the secondary labels among the top10.","metadata":{}},{"cell_type":"code","source":"# transfer secondary_labels from train_metadata to df_model_train\n\ndf_model_train = pd.merge(df_model_train, train_metadata[['secondary_labels','filename']], on='filename')\ndf_model_train['secondary_labels'] = df_model_train['secondary_labels'].apply(eval)\n\n# compute an intersection of top10 and secondary\n\ndf_model_train['secondary_in_top10'] = [set(a).intersection(set(b)) for a,b in zip(df_model_train['top10_labels'],df_model_train['secondary_labels'])]\n\n# summary\n\nprint('total records with secondary labels {}'.format( sum(df_model_train['secondary_labels'].apply(len)>0) ) )\n\nprint('model finds secondary among the top10 in {} cases'.format(sum(df_model_train['secondary_in_top10'].apply(len)>0) ) )","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:12.826822Z","iopub.execute_input":"2023-04-04T03:53:12.827187Z","iopub.status.idle":"2023-04-04T03:53:12.981063Z","shell.execute_reply.started":"2023-04-04T03:53:12.827151Z","shell.execute_reply":"2023-04-04T03:53:12.979614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we should account for one caveat. It is possible that all 10 samples in top10 have the primary label. The model could still be guessing the second-best unique label correctly, but we can't tell because of the top10 cutoff in our dataframe. \n\n\nLet's see in how many cases this could be true.","metadata":{}},{"cell_type":"code","source":"index_not_in_top10 = (df_model_train['secondary_labels'].apply(len)>0) * (df_model_train['secondary_in_top10'].apply(len)==0)\n\nprint(\"in {} cases, we don't know if model's second best label is among secondary labels, \"\n      \"because all top 10 labels are the same\"\n      .format(sum(df_model_train.loc[index_not_in_top10, 'top10_labels'].apply(set).apply(len)==1))\n     )\n","metadata":{"execution":{"iopub.status.busy":"2023-04-04T03:53:12.982825Z","iopub.execute_input":"2023-04-04T03:53:12.983934Z","iopub.status.idle":"2023-04-04T03:53:13.329680Z","shell.execute_reply.started":"2023-04-04T03:53:12.983884Z","shell.execute_reply":"2023-04-04T03:53:13.328554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Summary:** In _at least_ 1251 of 2305 cases, the model finds secondary labels correctly. ","metadata":{}},{"cell_type":"markdown","source":"### <font color='289C4E'>I hope you enjoyed reading this notebook and got new ideas for your own analysis! <font>","metadata":{}}]}