{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#### Thank you for looking at this notebook. I think you'll find the pictures and sounds insightful and interesting. Please upvote and good luck on the competition!\n\n#### ---Paul Alton Nussbaum","metadata":{}},{"cell_type":"markdown","source":"# Executive Summary\n* This notebook allows the user to modify parameters along the classification pipeline and observe the results. As with traditional observation methods, the notebook lets users view input vectors for similar and different birds.\n* In addition to traditional methods, this notebook also presents data in its original format (audio recordings of birds). This is intuitive for someone testing a microphone and recording system - they will want to listen to the recordings to see if they contain valid and useful information.\n* The notebook extends this intuitive and useful technique to individual neural network layers - working backwards towards a best estimate of the original input (a system I call \"reading the robot mind\").\n* The user can even provide just the \"answer\" (select a bird at the final output layer of the AI solution), and the reading the robot mind system will work backwards through the entire automated process and AI layers to let the user hear a best approximation of what it has learned that bird sounds like.\n* You may be surprised at what you hear!","metadata":{}},{"cell_type":"markdown","source":"# Introduction\nAs an AI practitioner, you know the value of being able to explain your solutions to subject matter experts in ways that they are familiar with. \n\n#### Even when that expert is you!\n\nIf you designed a complex autonomous system, you would want to be able to check the data through every step of the process. Your autonomous system would be so open and transparent - you could observe the data at any point along the way. Most importantly, you would want that data presented in a format that is familiar to you. \n\nFor example: If you are making an automated process that classifies birds, you'd want to test the microphones, the digital storage of the data, and the segmentation of the data stream to present for classification. You want to hear and possibly visualize the data quality every step of the way. In effect, you are taking a sequence of numbers stored on disk - an audio file segment - and recreating a best approximation of input (the sounds that the microphone sought to capture). If the recreated sound is garbled or otherwise doesn't contain enough information for you to hear the difference between different birds, you have a chance to correct the problem before moving on to the next step in your automated process. \n\n####  Don't you want a system to accomplish this for each layer of your AI?\n\nI call this system \"reading the robot mind.\"\n\nIt can be easily accomplished by simultaneously training auto-encoders along each layer of the network (able to recreate the best approximation of the original input from the data available at that stage of the process). Many shy from this approach because of the added computational cost and time - as competition forces a race to market. I think it is a mistake to skip this step because the training cost is minimal compared to inference costs if the solution is successful. As an example, I cite the currently running Kagle competition where these autoencoders were not trained along with the generative solution, and now there is a scramble to \"read the robot mind,\" (https://www.kaggle.com/competitions/stable-diffusion-image-to-prompts) but there are many other examples, even with in-the-news stories about ChatGPT and the search for inputs that cause toxic or undesired responses.\n\nI have created this notebook to show that it is possible to make a system to read the robot mind, even when those autoencoders were not trained alongside the network.\n\nThe notebook lets you hear, visualize, and compare data every step and layer on the way. It allows comparison of samples from the same type of bird (so you can see and hear if they are similar) and different birds (so you can see and hear if they are different) including:\n* Segmentation\n* Feature extraction (choice of Mel Spectrogram or MFCC, both having user selectable parameters)\n* Conversion to image format for archival and training, including reduction to 8bit quantization.\n* Each layer of a Conv2d, MaxPool combo model, followed by Dense embedding and Dense softmax categorizer.\n\nThis notebook provides the ability to work backwards through the system from any point and let the user see a best estimation of the features that were presented as input, as well as hear a best approximation of the bird recording from which they were extracted.","metadata":{}},{"cell_type":"markdown","source":"* v15 - Add some more explanatory documentation for publication, fine tune\n* v14 - Allow working backwards from a manually selected classification \"answer\" back to a best approximation of the bird sound associated with that classification.\n* v13 - Improve visualization\n* v12 - Include dense layers\n* v11 - include audio playback of deep layer reconstructions.\n* v10 - Do not save image files if already saved, but allow tuning of feature extraction algorithms. Also visuallization into deep NN (multi-layer cumulative filters as well as visualizing input data as it makes its way through each layer of the network.\n* v09 - Cleaning up code. Breaking out image saving, so it can be used in a different notebook. Also submission (depending on time required).\n* v08 - Skipping segmentation adjustments (stick to 5 sec segments). Allow choice of Mel Spectrum or MFCC, and then choice of Mel Spectrum Band Counts and (if using MFCC) Coefficient counts.\n* v07 - Explaining the parts of \"Reading the Robot Mind\"\n* v06 - Experimenting with fidelity of transformations and reverse transformations.\n* v05 - converting audio files into 5 second MFCC graphics images\n* v01-v04 - starting from scratch to set up \"reading the robot mind\" examples.","metadata":{}},{"cell_type":"markdown","source":"## Loading the meta data","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ntrain_meta = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\ntotal_data_count = len(train_meta)\nprint(\"There are a total of \", total_data_count, \"  entries in the dataset\")\n# Uncomment the below two lines if you wish to see an example nmeta data entry.\n# print(\"Here is the first entry in the dataset metadata:\\n\")\n# print(train_meta.iloc[0])\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:33:53.611657Z","iopub.execute_input":"2023-04-24T17:33:53.612456Z","iopub.status.idle":"2023-04-24T17:33:53.803602Z","shell.execute_reply.started":"2023-04-24T17:33:53.612392Z","shell.execute_reply":"2023-04-24T17:33:53.80085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## IMPORTANT - If you are re-running this notebook, use only a few different bird types to speed up your experimentation\nYou can then switch to \"all birds\" once things are working to your satisfaction. See code block below.","metadata":{}},{"cell_type":"code","source":"# Create a list of all the possible classifications (labels)\nlabels = list(train_meta['primary_label'].unique())\nprint(\"There are \", len(labels), \" different labels in the whole dataset.\\n\")\nnum_birds = len(labels)\nepochs = 2\n\n# Limit the number of birds to speed things up for the user\n# Remove the next lines if using the entire set of birds\n# num_birds = 9 # was 9\n# labels = labels[0:num_birds]\n# print(\"To speed things up, we are only using \", len(labels), \" different labels (or birds).\\n\")\n# epochs = 5 ","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:37:55.801231Z","iopub.execute_input":"2023-04-24T17:37:55.801965Z","iopub.status.idle":"2023-04-24T17:37:55.816009Z","shell.execute_reply.started":"2023-04-24T17:37:55.801905Z","shell.execute_reply":"2023-04-24T17:37:55.814685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The notebook will convert audio into images, the result of segmentation and feature extraction","metadata":{}},{"cell_type":"code","source":"image_data = []\nfor label in labels : \n    image_data.append(train_meta[train_meta[\"primary_label\"] == label].iloc[:])\nimage_data = pd.concat(image_data).reset_index(drop=True) \ntotal_image_data_count = len(image_data)\nprint(\"There are \", total_image_data_count, \" different entries in the image data set.\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:33:53.828649Z","iopub.execute_input":"2023-04-24T17:33:53.829002Z","iopub.status.idle":"2023-04-24T17:33:53.862574Z","shell.execute_reply.started":"2023-04-24T17:33:53.828966Z","shell.execute_reply":"2023-04-24T17:33:53.860982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# un-comment if you wish to see the meta data\n#image_data","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:33:53.865722Z","iopub.execute_input":"2023-04-24T17:33:53.866409Z","iopub.status.idle":"2023-04-24T17:33:53.871844Z","shell.execute_reply.started":"2023-04-24T17:33:53.866372Z","shell.execute_reply":"2023-04-24T17:33:53.870474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Helper functions to transform and untransform (reverse transform) and save resultant features","metadata":{}},{"cell_type":"code","source":"# These functions will help transform and untransform (reverse transformation) as well as\n# reducing the resolution of floating point feature values into 8-bit unsigned integers.\n# The goal is to convert each 5 second segment of auio into a 2D greyscale image of 8-bit uints.\n# After this is done, we can use these images as inputs to our classifier.\n\nimport librosa \nimport numpy as np\nfrom PIL import Image\nimport os\nimport soundfile as sf\n# from tqdm import tqdm\n\n\n# Convert 1D audio amplitude over time, to another domain (2 dimensional \"image\")\n#####################################################\n###                                               ###\n### ALLOW USER TO MODIFY THESE AND SEE THE IMPACT ###\n###                                               ###\n#####################################################\nuse_mfcc = False # otherwise use Mel Spectrogram\nn_mels = 64 # used for both Mel Spectrogram and MFCC\nn_coeff = 64 # number of cepstral coefficients (only used for MFCC)\n\n\n# Image dimensions\nif use_mfcc :\n    if n_coeff < n_mels :\n        num_rows = n_coeff \n    else :\n        num_rows = n_mels\nelse :\n    num_rows = n_mels\nnum_columns = 313 # number of MFCC frames generated from 160000 samples of sudio (5 sec at 32000)\nnum_channels = 1\n\n# MAY BE A MISTAKE TO MAKE THIS A GLOBAL VALUE - but it's always forced to 32000\nsr = 32000 # default sampling rate if none provided\n\nWindow_Size = 5 # 5 seconds\n# Get number of samples for 5 seconds\nsegment = Window_Size * sr\n\n\ndef Audio_to_Domain(y, sr):\n    if use_mfcc:\n        feat = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=n_coeff, n_mels=n_mels, power=3) # raise to 3rd power before extracting coefficients to reduce noise\n    else :\n        feat = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=n_mels, power=1)\n    if feat.shape[1] <= num_columns:\n        pad_width = num_columns - feat.shape[1]\n        feat = np.pad(feat, pad_width=((0,0),(0,pad_width)), mode='constant')\n    return feat\n\ndef Domain_to_Audio(feat, sr):\n    if use_mfcc :\n        y = librosa.feature.inverse.mfcc_to_audio(feat, sr=sr, n_mels=n_mels, power=3)\n    else :\n        y = librosa.feature.inverse.mel_to_audio(feat, sr=sr, power=1) \n    return y\n\n\ndef Audio_Segment_Transform_and_Save(file_path, cat_name='dummy', is_save=True, train=True, sr=sr):\n\n    # First load the file\n    filename = file_path.replace(\"/\", \"_\")\n    file_path = \"/kaggle/input/birdclef-2023/train_audio/\" + file_path\n    audio, sr = librosa.load(file_path, sr = sr)\n\n    # Get number of samples for 5 seconds\n    buffer = segment\n\n    samples_total = len(audio)\n    samples_wrote = 0\n    counter = 1\n\n    # this will be a list of 5 second segments\n    feature_split = []\n    # each segment will have its own filename\n    feature_filenames = []\n    # each segment will also have its own min and range value for reconstruction (reverse transformation)\n    feature_mins = []\n    feature_ranges = []\n    while samples_wrote < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_wrote):\n            buffer = samples_total - samples_wrote\n\n        block = audio[samples_wrote : (samples_wrote + buffer)]\n        these_features = Audio_to_Domain(block, sr)\n        feature_split.append(these_features)\n\n        # Write Window_Size second segment\n        if is_save == True:\n            # had to add the letter \"z\" in front of my output files and paths, becuase contest submission searches output files alphabetically (!?)\n            out_filename = \"/kaggle/working/zimages/\" + cat_name + \\\n                \"/split_\" + str(counter) + \"_\" + filename + \".png\"\n            \n            feature_filenames.append(out_filename)\n            \n            # convert the features into an image for saving\n            \n            #scale to 0 - 255\n            fmax = these_features.max()\n            fmin = these_features.min()\n            frange = fmax - fmin\n            # these_features = 255 * ((these_features - fmin)/frange)\n            these_features = np.array((((these_features - fmin) / frange)*255), dtype='uint8')\n            # save the min and range for later reconstruction (reverse transform, or un-transform) back to audio\n            feature_mins.append(fmin)\n            feature_ranges.append(frange)\n            this_image = Image.fromarray(these_features, mode = 'L')\n            # make the directory if it doesn't yet exist\n            if not os.path.exists(os.path.dirname(out_filename)):\n                try:\n                    os.makedirs(os.path.dirname(out_filename))\n                except OSError as exc: # Guard against race condition\n                    if exc.errno != errno.EEXIST:\n                        raise            \n            # save the file\n            this_image.save(out_filename, format=\"PNG\")\n            \n        counter += 1\n        samples_wrote += buffer\n    return feature_split, sr, feature_filenames, feature_mins, feature_ranges\n\ndef Load_and_UnTransform(filename):\n    this_image = Image.open(filename)\n    block = Domain_to_Audio(this_image, sr)\n    return block, sr\n    \n\ndef Save_Features(_df, train=True):\n    data = []\n\n    if os.path.exists(os.path.dirname(\"/kaggle/working/zimages/\")):\n        print(\"WARNING: Image files already saved. Aborting save to avoid corrupted image databse.\")\n        data_df = pd.DataFrame(data, columns=['primary_label', 'original_filename', 'filename', 'fmin', 'frange'])\n        # had to add the letter z in front of filenames and paths so competition submission can find the output file (alphabetically)\n        data_df = pd.read_csv(\"/kaggle/working/zimages.csv\")\n    else:\n        for index, row in _df.iterrows(): # was tqdm(_df.iterrows()):\n            audio_lst, sr, filenames, mins, ranges = Audio_Segment_Transform_and_Save(row[\"filename\"], cat_name = row[\"primary_label\"], is_save = True, train = train)\n            for idx, y in enumerate(audio_lst):\n                data.append([row[\"primary_label\"], row[\"filename\"], filenames[idx], mins[idx], ranges[idx]])\n    \n        data_df = pd.DataFrame(data, columns=['primary_label', 'original_filename', 'filename', 'fmin', 'frange'])\n        data_df.to_csv(\"/kaggle/working/zimages.csv\", index=False)\n\n    return data_df\n    \n    \n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:33:53.874615Z","iopub.execute_input":"2023-04-24T17:33:53.875354Z","iopub.status.idle":"2023-04-24T17:33:53.964412Z","shell.execute_reply.started":"2023-04-24T17:33:53.875051Z","shell.execute_reply":"2023-04-24T17:33:53.962497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# See and Hear the Feature Extraction\n#### Try a number of Feature Extraction Algorithms and test the following:\n* Do the features visually look similar for the same bird, and different for different birds?\n* Is the similarity/difference enough to be able to visually classify which bird is which?\n* What about the audio - is the recreated sound clear enough for an expert to identify the bird?\n* If the answer is \"no\" to any of the above, allow fine tuning by the user (three different feature extractions are compared below)","metadata":{}},{"cell_type":"code","source":"# Visually examine feature extraction algorithms\nimport IPython\nimport matplotlib.pyplot as plt\n\ndef Compare_Feature_Extraction(image_data, list_to_compare) :\n    global sr, segment\n    list_len = len(list_to_compare)\n    if list_len < 2 :\n        print(\"Must provide at least 2 items - cannot compare.\\n\")\n        return()\n    else:\n        plt.figure(figsize=(10, 10))\n        for i in range(list_len):\n            # grab from the image set\n            audio_filename = image_data.at[list_to_compare[i],\"filename\"]\n            common_name = image_data.at[list_to_compare[i],\"common_name\"]\n            # load the audio data\n            audio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n            # Take first 5 second \"segment\"\n            audio = audio[0:segment]\n            # Extract the feaures\n            feat = Audio_to_Domain(audio, sr)\n            # Move the range from current min and max, into 0 to 255 8 bit integers\n            fmin = feat.min()\n            fmax = feat.max()\n            frange = fmax - fmin\n            feat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\n            # Plot the features\n            ax = plt.subplot(3, 3, i + 1)\n            plt.imshow(feat, cmap='gray', aspect=(num_columns/num_rows), interpolation = 'None')\n#            plt.imshow(images[i].numpy().astype(\"uint8\"), cmap='gray', aspect=(num_columns/num_rows), interpolation = 'None')\n            plt.title(common_name)\n            plt.axis(\"off\")\n    return()\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:33:53.966632Z","iopub.execute_input":"2023-04-24T17:33:53.967011Z","iopub.status.idle":"2023-04-24T17:33:53.980338Z","shell.execute_reply.started":"2023-04-24T17:33:53.966973Z","shell.execute_reply":"2023-04-24T17:33:53.978512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's start with Mel Scale Frequency Spectrum feature extraction, with a resultant 16 mel scale bands","metadata":{}},{"cell_type":"code","source":"#####################################################\n###                                               ###\n### ALLOW USER TO MODIFY THESE AND SEE THE IMPACT ###\n###                                               ###\n#####################################################\nuse_mfcc = False # otherwise use Mel Spectrogram\nn_mels = 16 # used for both Mel Spectrogram and MFCC\nn_coeff = 1 # number of cepstral coefficients (only used for MFCC)\n\n# These are calculated based on user selections\n# Image dimensions\nif use_mfcc :\n    if n_coeff < n_mels :\n        num_rows = n_coeff \n    else :\n        num_rows = n_mels\nelse :\n    num_rows = n_mels","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:33:53.982254Z","iopub.execute_input":"2023-04-24T17:33:53.982748Z","iopub.status.idle":"2023-04-24T17:33:53.994274Z","shell.execute_reply.started":"2023-04-24T17:33:53.982712Z","shell.execute_reply":"2023-04-24T17:33:53.992885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Compare features from the SAME birds (first 5 sec of audio used)\\n\")            \n# Compare the first few entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 1, 2, 3, 4, 5, 6, 7, 8])","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:33:53.995803Z","iopub.execute_input":"2023-04-24T17:33:53.996549Z","iopub.status.idle":"2023-04-24T17:34:11.197896Z","shell.execute_reply.started":"2023-04-24T17:33:53.996494Z","shell.execute_reply":"2023-04-24T17:34:11.196504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from the same bird\n* Comparing features from the same bird yields some similarities, but also differences (maybe different songs of the same bird)\n* Notice how the third from last (lower left) seems to have not much of a signal (we are not changing segmentation algorithms, but maybe there is a problem with segmentation in that the 5 sec window starting from the beginning of audio, might not be an optimal window).\n* Maybe we are throwing away too much information.","metadata":{}},{"cell_type":"code","source":"print(\"Compare features from the DIFFERENT birds (first 5 sec of audio used)\\n\")            \n# Compare the first 10 entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 50, 150, 200, 225, 350, 430])","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:11.200061Z","iopub.execute_input":"2023-04-24T17:34:11.201155Z","iopub.status.idle":"2023-04-24T17:34:12.46006Z","shell.execute_reply.started":"2023-04-24T17:34:11.201099Z","shell.execute_reply":"2023-04-24T17:34:12.458774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from DIFFERENT birds\n* Comparing features from the different birds yields sufficient differences","metadata":{}},{"cell_type":"markdown","source":"## Finally, let's listen to the untransformation (recreate audio from the extracted features)\n* We'll only listen to the first segment from the first bird","metadata":{}},{"cell_type":"code","source":"audio_filename = image_data.at[0,\"filename\"]\ncommon_name = image_data.at[0,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n# Extract the feaures\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\n# Go back to the original range of values\nfeat = ((feat/255)*frange)+fmin\n# Recreate the audio from the extracted features\nrecreated_audio = Domain_to_Audio(feat, sr)\n\nprint(\"In BLUE is the original audio of the first segment of the first audio file\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"Superimposed in ORANGE the RECREATED (un-transformed) audio of the first segment of the first image file\\n\")\nlibrosa.display.waveshow(recreated_audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:12.464518Z","iopub.execute_input":"2023-04-24T17:34:12.464904Z","iopub.status.idle":"2023-04-24T17:34:17.041059Z","shell.execute_reply.started":"2023-04-24T17:34:12.464867Z","shell.execute_reply":"2023-04-24T17:34:17.039468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"This is the audio playback of the RECREATED audio file\")\nIPython.display.Audio(data = recreated_audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:17.043291Z","iopub.execute_input":"2023-04-24T17:34:17.043831Z","iopub.status.idle":"2023-04-24T17:34:17.064985Z","shell.execute_reply.started":"2023-04-24T17:34:17.04378Z","shell.execute_reply":"2023-04-24T17:34:17.063612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above lets us listen to recreation of the original input from the Extracted Features\n* The audio is a bit garbled but still clearly a bird call\n* The real question is \"will a human expert be able to use the RECREATED audio to correctly classify birds.\"\n* The answer will be a qualitative measure of the feature extraction algorithm","metadata":{}},{"cell_type":"markdown","source":"## Now let's try another Feature Extraction Algorithm\n### This time let's try MFCC with 16 coefficients","metadata":{}},{"cell_type":"code","source":"# Convert 1D audio amplitude over time, to another domain (2 dimensional \"image\")\n#####################################################\n###                                               ###\n### ALLOW USER TO MODIFY THESE AND SEE THE IMPACT ###\n###                                               ###\n#####################################################\nuse_mfcc = True # otherwise use Mel Spectrogram\nn_mels = 64 # used for both Mel Spectrogram and MFCC\nn_coeff = 16 # number of cepstral coefficients (only used for MFCC)\n\n\n# Image dimensions\nif use_mfcc :\n    if n_coeff < n_mels :\n        num_rows = n_coeff \n    else :\n        num_rows = n_mels\nelse :\n    num_rows = n_mels","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:17.066551Z","iopub.execute_input":"2023-04-24T17:34:17.068727Z","iopub.status.idle":"2023-04-24T17:34:17.092782Z","shell.execute_reply.started":"2023-04-24T17:34:17.068681Z","shell.execute_reply":"2023-04-24T17:34:17.091347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Compare features from the SAME birds (first 5 sec of audio used)\\n\")            \n# Compare the first 10 entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 1, 2, 3, 4, 5, 6, 7, 8])","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:17.094569Z","iopub.execute_input":"2023-04-24T17:34:17.095024Z","iopub.status.idle":"2023-04-24T17:34:18.634726Z","shell.execute_reply.started":"2023-04-24T17:34:17.094976Z","shell.execute_reply":"2023-04-24T17:34:18.633252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from the same bird### The above showed...\n* Comparing features from the same bird yields fewer similarities, and many differences\n* Maybe we are throwing away too much information.","metadata":{}},{"cell_type":"code","source":"print(\"Compare features from the DIFFERENT birds (first 5 sec of audio used)\\n\")            \n# Compare the first 10 entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 50, 150, 200, 225, 350, 430])","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:18.636523Z","iopub.execute_input":"2023-04-24T17:34:18.636922Z","iopub.status.idle":"2023-04-24T17:34:19.923999Z","shell.execute_reply.started":"2023-04-24T17:34:18.636885Z","shell.execute_reply":"2023-04-24T17:34:19.922343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from DIFFERENT birds\n* Different birds indeed seem to have different features, but also striking similarities (Emerald Cuckoo and Fish-Eagle)\n* Not sure if \"different versus similar\" birds really jump out at me personally when I view these\n* Not sure if a bird expert would use these images to classify birds.","metadata":{}},{"cell_type":"markdown","source":"## Let's listen to the untransformation (recreate audio from the extracted features)\n* We'll only listen to the first segment from the first bird","metadata":{}},{"cell_type":"code","source":"audio_filename = image_data.at[0,\"filename\"]\ncommon_name = image_data.at[0,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n# Extract the feaures\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\n# Go back to the original range of values\nfeat = ((feat/255)*frange)+fmin\n# Recreate the audio from the extracted features\nrecreated_audio = Domain_to_Audio(feat, sr)\n\nprint(\"In BLUE is the original audio of the first segment of the first audio file\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"Superimposed in ORANGE the RECREATED (un-transformed) audio of the first segment of the first image file\\n\")\nlibrosa.display.waveshow(recreated_audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:19.926016Z","iopub.execute_input":"2023-04-24T17:34:19.926396Z","iopub.status.idle":"2023-04-24T17:34:23.439783Z","shell.execute_reply.started":"2023-04-24T17:34:19.926361Z","shell.execute_reply":"2023-04-24T17:34:23.438218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"This is the audio playback of the RECREATED audio file\")\nIPython.display.Audio(data = recreated_audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:23.441562Z","iopub.execute_input":"2023-04-24T17:34:23.441936Z","iopub.status.idle":"2023-04-24T17:34:23.459932Z","shell.execute_reply.started":"2023-04-24T17:34:23.441901Z","shell.execute_reply":"2023-04-24T17:34:23.458477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above lets us listen to recreation of the original input from the Extracted Features## The above showed\n* The audio is a bit garbled but better than just the 16 band Mel Scale Spectrum - still clearly a bird call\n* The real question is \"will a human expert be able to use the RECREATED audio to correctly classify birds.\"\n* The answer will be a qualitative measure of the feature extraction algorithm\n* From my non-expert listening, I'd say these 16 parameters are better than the prior 16 parameters (audio wise)","metadata":{}},{"cell_type":"markdown","source":"## Finally; let's finish with a 32 Mel Scale frequency band feature extraction","metadata":{}},{"cell_type":"code","source":"# Convert 1D audio amplitude over time, to another domain (2 dimensional \"image\")\n#####################################################\n###                                               ###\n### ALLOW USER TO MODIFY THESE AND SEE THE IMPACT ###\n###                                               ###\n#####################################################\nuse_mfcc = False # otherwise use Mel Spectrogram\nn_mels = 32 # used for both Mel Spectrogram and MFCC\nn_coeff = 1 # number of cepstral coefficients (only used for MFCC)\n\n\n# Image dimensions\nif use_mfcc :\n    if n_coeff < n_mels :\n        num_rows = n_coeff \n    else :\n        num_rows = n_mels\nelse :\n    num_rows = n_mels","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:23.461828Z","iopub.execute_input":"2023-04-24T17:34:23.462168Z","iopub.status.idle":"2023-04-24T17:34:23.468199Z","shell.execute_reply.started":"2023-04-24T17:34:23.462126Z","shell.execute_reply":"2023-04-24T17:34:23.467282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Compare features from the SAME birds (first 5 sec of audio used)\\n\")            \n# Compare the first 10 entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 1, 2, 3, 4, 5, 6, 7, 8])","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:23.470313Z","iopub.execute_input":"2023-04-24T17:34:23.470898Z","iopub.status.idle":"2023-04-24T17:34:24.882313Z","shell.execute_reply.started":"2023-04-24T17:34:23.470845Z","shell.execute_reply":"2023-04-24T17:34:24.880954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from the same bird\n* Comparing features from the same bird now yields more similarities\n* Seems like now we have plenty of information (maybe too much? Will it slow down our training algorithm or cause it to overfit?)","metadata":{}},{"cell_type":"code","source":"print(\"Compare features from the DIFFERENT birds (first 5 sec of audio used)\\n\")            \n# Compare the first 10 entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 50, 150, 200, 225, 350, 430])","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:24.884136Z","iopub.execute_input":"2023-04-24T17:34:24.885966Z","iopub.status.idle":"2023-04-24T17:34:26.118402Z","shell.execute_reply.started":"2023-04-24T17:34:24.885911Z","shell.execute_reply":"2023-04-24T17:34:26.117115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from DIFFERENT birds\n* Different birds indeed seem to have different features, but also instantaneous similarities (which may be differentiated as they change over time). \n* \"Different versus similar\" birds start to make itself more visible, but with regard to that time window, it looks like 1.5 to 2 seconds of changing notes might be needed to differentiate birds (since there are 313 samples in a 5 second window, this equates to 95 to 125 feature vectors.\n* Still not sure if a bird expert would use these images to classify birds.","metadata":{}},{"cell_type":"markdown","source":"## Let's listen to the untransformation (recreate audio from the extracted features)\n* We'll only listen to the first segment from the first bird","metadata":{}},{"cell_type":"code","source":"audio_filename = image_data.at[0,\"filename\"]\ncommon_name = image_data.at[0,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n# Extract the feaures\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\n# Go back to the original range of values\nfeat = ((feat/255)*frange)+fmin\n# Recreate the audio from the extracted features\nrecreated_audio = Domain_to_Audio(feat, sr)\n\nprint(\"In BLUE is the original audio of the first segment of the first audio file\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"Superimposed in ORANGE the RECREATED (un-transformed) audio of the first segment of the first image file\\n\")\nlibrosa.display.waveshow(recreated_audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:26.122722Z","iopub.execute_input":"2023-04-24T17:34:26.123083Z","iopub.status.idle":"2023-04-24T17:34:29.63088Z","shell.execute_reply.started":"2023-04-24T17:34:26.123047Z","shell.execute_reply":"2023-04-24T17:34:29.629648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"This is the audio playback of the RECREATED audio file\")\nIPython.display.Audio(data = recreated_audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:29.63254Z","iopub.execute_input":"2023-04-24T17:34:29.63376Z","iopub.status.idle":"2023-04-24T17:34:29.653021Z","shell.execute_reply.started":"2023-04-24T17:34:29.63372Z","shell.execute_reply":"2023-04-24T17:34:29.649546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above lets us listen to recreation of the original input from the Extracted Features## The above showed\n* The audio is definitely more clear\n* It seems obvious that \"a human expert will be able to use the RECREATED audio to correctly classify birds.\"\n* Naturally, if they couuld perform the classification with the raw audio\n* Segmentation is still the open question - if the 5 second window does not containt enough of that bird's call, etc...","metadata":{}},{"cell_type":"markdown","source":"# Save Training Data\n## For each raw audio file, we will use the selected Feature Extraction to convert the audio into feature files to be used for training\n* Segment using non-verlapping 5 minute windows (tail-end windows may be shorter, ignore FFT warnings and pad with 0's)\n* Perform the selected feature extraction, converting the 1D audio into the 2D image (pixels are 8-bit resolution)\n* Save these images to be used as a dataset for training\n* Also save the scaling factors for each image in the data set (so we can use it to recreate audio). Scaling factor is not used by the NN.","metadata":{}},{"cell_type":"code","source":"import time \n\ndef Save_Image_Examples():\n    global image_df, n_mels, n_coeff, use_mfcc, num_rows\n    # Image dimensions\n    if use_mfcc :\n        if n_coeff < n_mels :\n            num_rows = n_coeff \n        else :\n            num_rows = n_mels\n    else :\n        num_rows = n_mels\n\n    print(\"Starting to save Images. Please wait...\\n\")\n    start_time = time.time()\n    image_df = Save_Features(image_data, train=True)\n    end_time = time.time()\n    print(\"Total amount of time to convert and save image data is \", (end_time - start_time), \n          \" seconds\\n\")\n    return","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:29.654795Z","iopub.execute_input":"2023-04-24T17:34:29.65517Z","iopub.status.idle":"2023-04-24T17:34:29.665203Z","shell.execute_reply.started":"2023-04-24T17:34:29.655133Z","shell.execute_reply":"2023-04-24T17:34:29.663769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"There are \", total_image_data_count, \" different entries in the audio files data set.\\n\")\nSave_Image_Examples()\n#Save_Validation_Examples()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:34:29.667878Z","iopub.execute_input":"2023-04-24T17:34:29.66831Z","iopub.status.idle":"2023-04-24T17:36:46.204762Z","shell.execute_reply.started":"2023-04-24T17:34:29.668273Z","shell.execute_reply":"2023-04-24T17:36:46.203376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test Saved Training Data\n## Let's make sure everything got saved properly\n* Read in one of the saved image files\n* View the features (as an image file)\n* Untransform the extracted features so we can see and hear the reconstructed audio","metadata":{}},{"cell_type":"code","source":"import PIL\n# Let's look and at the first validation transformed image data \n# Read in the first validation segment graphic\ntransformed_file = image_df.at[0,\"filename\"]\noriginal_file = image_df.at[0,\"original_filename\"]\n\ntrans_1 = PIL.Image.open(transformed_file)\nprint(\"Shape of loaded image  \", trans_1.height, trans_1.width)\naudio_1, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + original_file, sr = sr)\n# Let's look only at the first segment (of size segment = 5 sec * sampling rate)\naudio_1 = audio_1[0:segment]","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:46.206969Z","iopub.execute_input":"2023-04-24T17:36:46.207822Z","iopub.status.idle":"2023-04-24T17:36:46.292487Z","shell.execute_reply.started":"2023-04-24T17:36:46.207773Z","shell.execute_reply":"2023-04-24T17:36:46.290999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"This is the original audio, (segmented but not transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(audio_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:46.29424Z","iopub.execute_input":"2023-04-24T17:36:46.29553Z","iopub.status.idle":"2023-04-24T17:36:46.757514Z","shell.execute_reply.started":"2023-04-24T17:36:46.295474Z","shell.execute_reply":"2023-04-24T17:36:46.755842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import IPython\n\nIPython.display.Audio(data = audio_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:46.75926Z","iopub.execute_input":"2023-04-24T17:36:46.759629Z","iopub.status.idle":"2023-04-24T17:36:46.776206Z","shell.execute_reply.started":"2023-04-24T17:36:46.759593Z","shell.execute_reply":"2023-04-24T17:36:46.7747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of loaded image (height and width)\", trans_1.height, trans_1.width)\nprint(\"Graphic of loaded image file\")\nprint(\"Here is the image file that was loaded\\n\")\nplt.imshow(trans_1, cmap='gray', aspect=int(num_columns/num_rows) , interpolation = 'None')\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:46.778075Z","iopub.execute_input":"2023-04-24T17:36:46.778453Z","iopub.status.idle":"2023-04-24T17:36:47.004181Z","shell.execute_reply.started":"2023-04-24T17:36:46.778398Z","shell.execute_reply":"2023-04-24T17:36:47.002624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nnorm_trans = np.array(trans_1)\nprint(\"shape of image array:\", norm_trans.shape)\nprint(\"Before scaling, image has values within this range: \", norm_trans.min(), norm_trans.max())\n# re-scale the transformation back to its original min, max, and range, using the values we saved earlier\nrescale_factor = image_df.at[0,\"frange\"]/255\nrescale_shift = image_df.at[0,\"fmin\"]\nnorm_trans = (norm_trans * rescale_factor) + rescale_shift\n\nprint(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:47.012967Z","iopub.execute_input":"2023-04-24T17:36:47.013486Z","iopub.status.idle":"2023-04-24T17:36:50.092716Z","shell.execute_reply.started":"2023-04-24T17:36:47.013424Z","shell.execute_reply":"2023-04-24T17:36:50.091064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:50.094578Z","iopub.execute_input":"2023-04-24T17:36:50.095253Z","iopub.status.idle":"2023-04-24T17:36:50.112555Z","shell.execute_reply.started":"2023-04-24T17:36:50.095215Z","shell.execute_reply":"2023-04-24T17:36:50.111616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above shows...\n* We have successfully created a new dataset, of image files, instead of audio files\n* We have confidence about the information content of the saved data containing the extracted features\n* * This is qualitatively shown using the \"Reading the Robot Mind\" method of seeing if a human expert can classify the untransformed audio","metadata":{}},{"cell_type":"markdown","source":"# Build Convolutional Neural Network\n## Build a simple CNN to perform classification\n* Use sequences of Conv2D and MaxPool to find edges, textures, and patterns\n* NOTE: You can change Conv2D filter counts and shapes, as well as MaxPool2D shapes, but if you add or remove layers, you'll need to modify the backwards process (manually coded later in the notebook) as I have not yet implemented a method that figures out the layer structure automatically.","metadata":{}},{"cell_type":"code","source":"from tensorflow import keras\nfrom keras.models import Model\nfrom keras.layers import Dense,Conv2D,Flatten,MaxPool2D,Concatenate, Activation\nfrom keras.layers import Dropout,BatchNormalization, Input, Rescaling\n\ninputs = Input(shape = (num_rows, num_columns, 1))\n\n# Rescale the input (arrives as 0-255 8-bit uint grayscale pixels)\nmodel = Rescaling(1./255)(inputs)\n\n# This model has a sequential set of Conv2D and MaxPool layers finished with a 32 Dense \n# In this way, the model seeks to embed the entire image into a 32 dim vector\nmodel1 = Conv2D(filters=16, kernel_size=(3, 3), padding='SAME', activation='relu')(model)   # was 32 filters\nmodel1 = Conv2D(filters=16, kernel_size=(3, 3), padding='SAME', activation='relu')(model1)  # was 32 filters\nmodel1 = MaxPool2D(pool_size=(2, 2))(model1)\n\nmodel1a = Conv2D(filters=32, kernel_size=(5, 5), padding='SAME', activation='relu')(model1) # was 64 filters\nmodel1a = Dropout(0.2)(model1a)\nmodel1a = Conv2D(filters=16, kernel_size=(3, 3), padding='SAME', activation='relu')(model1a)# was 32 filters\nmodel1a = Dropout(0.2)(model1a)\n\nmodel1b = Conv2D(filters=16, kernel_size=(3, 3), padding='SAME', activation='relu')(model1a)# was 32 filters\nmodel1b = Conv2D(filters=16, kernel_size=(3, 3), padding='SAME', activation='relu')(model1b)# was 32 filters\n# Lets compress the time down\nmodel1b = MaxPool2D(pool_size=(2, 2))(model1b)\n\nmodel1c = Conv2D(filters=32, kernel_size=(5, 5), padding='SAME', activation='relu')(model1b)# was 64 filters\n\n# Flatten and use Dense to find a lower dim embedding\nmodel1c = Flatten()(model1c)\nmodel1c = Dense(64, activation = \"relu\")(model1c)\n\nbird_classification = Dense(num_birds, activation = 'softmax')(model1c)\n\nmodel = Model(inputs=inputs, outputs=bird_classification)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:50.114233Z","iopub.execute_input":"2023-04-24T17:36:50.114867Z","iopub.status.idle":"2023-04-24T17:36:59.414247Z","shell.execute_reply.started":"2023-04-24T17:36:50.114829Z","shell.execute_reply":"2023-04-24T17:36:59.412722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Here you can see the Model Summary","metadata":{}},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:59.416178Z","iopub.execute_input":"2023-04-24T17:36:59.416566Z","iopub.status.idle":"2023-04-24T17:36:59.473939Z","shell.execute_reply.started":"2023-04-24T17:36:59.416529Z","shell.execute_reply":"2023-04-24T17:36:59.472717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Here you can see the architecture of the model","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\ntf.keras.utils.plot_model(model)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:59.475546Z","iopub.execute_input":"2023-04-24T17:36:59.475869Z","iopub.status.idle":"2023-04-24T17:36:59.676975Z","shell.execute_reply.started":"2023-04-24T17:36:59.475835Z","shell.execute_reply":"2023-04-24T17:36:59.675761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Setup Training Pipeline\n## Use a pipeline for training with \"image_dataset_from_directory\" for simplicity","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport tensorflow_datasets as tfds\n\nbatch_size = 256 # was 32 # num_birds\n\nimg_height = num_rows\nimg_width = num_columns\n\nprint(\"Use dataset_from_directory to gather training images.\\n\")\ntrain_ds = tf.keras.utils.image_dataset_from_directory(\n  \"/kaggle/working/zimages/\",\n  validation_split=0.2,\n  subset=\"training\",\n  seed=111,    \n  image_size=(img_height, img_width),\n  color_mode = \"grayscale\",\n  label_mode = 'categorical',\n  batch_size=batch_size)\n\nprint(\"\\nClass names for training are: \", train_ds.class_names)\n\nprint(\"\\nUse dataset_from_directory to gather validation images.\\n\")\nval_ds = tf.keras.utils.image_dataset_from_directory(\n  \"/kaggle/working/zimages/\",\n  validation_split=0.2,\n  subset=\"validation\",\n  seed=111,\n  image_size=(img_height, img_width),\n  color_mode = \"grayscale\",\n  label_mode = 'categorical',\n  batch_size=batch_size)\nprint(\"\\nClass names for validation are: \", val_ds.class_names)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:36:59.678621Z","iopub.execute_input":"2023-04-24T17:36:59.678976Z","iopub.status.idle":"2023-04-24T17:37:01.875482Z","shell.execute_reply.started":"2023-04-24T17:36:59.678938Z","shell.execute_reply":"2023-04-24T17:37:01.873963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the data\nHere are the first few images from the training dataset.","metadata":{}},{"cell_type":"markdown","source":"# Test Training Pipeline\n## Double check that the saved images are correctly read by the pipeline","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nclass_names = train_ds.class_names\n\nplt.figure(figsize=(10, 10))\nfor images, labels in train_ds.take(1):\n  for i in range(5):\n    ax = plt.subplot(3, 3, i + 1)\n    plt.imshow(images[i].numpy().astype(\"uint8\"), cmap='gray', aspect=(num_columns/num_rows), interpolation = 'None')\n    plt.title(class_names[i])\n    plt.axis(\"off\")","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:37:01.877673Z","iopub.execute_input":"2023-04-24T17:37:01.878159Z","iopub.status.idle":"2023-04-24T17:37:02.847066Z","shell.execute_reply.started":"2023-04-24T17:37:01.87811Z","shell.execute_reply":"2023-04-24T17:37:02.845608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train the CNN\n## Train the Neural Network over a few Epochs","metadata":{}},{"cell_type":"code","source":"AUTOTUNE = tf.data.AUTOTUNE\n\ntrain_ds = train_ds.cache().prefetch(buffer_size=AUTOTUNE)\nval_ds = val_ds.cache().prefetch(buffer_size=AUTOTUNE)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:37:02.84925Z","iopub.execute_input":"2023-04-24T17:37:02.850095Z","iopub.status.idle":"2023-04-24T17:37:02.868814Z","shell.execute_reply.started":"2023-04-24T17:37:02.85004Z","shell.execute_reply":"2023-04-24T17:37:02.867488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping\n# Had to add the letter \"z\" in front of filenames and paths written to output, because Conteste submission only lists the first few ouputs alphabetically\ncheckpoint_model_path = \"/kaggle/working/zclassifier\"\nmetric = \"val_accuracy\"\n\nmodel.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])\n\ncheckpointer = ModelCheckpoint(\n    filepath=checkpoint_model_path,\n    monitor=metric, verbose=1, save_best_only=True)\nes_callback = EarlyStopping(monitor=metric, patience=5, verbose=1)\n\nhistory = model.fit(\n    train_ds, \n    validation_data=val_ds,\n    callbacks=[checkpointer,es_callback],\n    verbose = 1, # was 1,\n    epochs = epochs\n)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:38:16.006092Z","iopub.execute_input":"2023-04-24T17:38:16.006524Z","iopub.status.idle":"2023-04-24T17:40:37.233255Z","shell.execute_reply.started":"2023-04-24T17:38:16.006489Z","shell.execute_reply":"2023-04-24T17:40:37.231309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make sure we can load the saved model\nmodel = keras.models.load_model(checkpoint_model_path)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:40:43.884893Z","iopub.execute_input":"2023-04-24T17:40:43.886611Z","iopub.status.idle":"2023-04-24T17:40:44.751856Z","shell.execute_reply.started":"2023-04-24T17:40:43.886555Z","shell.execute_reply":"2023-04-24T17:40:44.750105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Show the history of the training loss and accuracy (both for training dataset and validation dataset)","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef plot_his(history, _auto=False):\n    plt.figure(1, figsize = (15,8))\n    plt.subplot(221)\n    if (_auto) :\n        plt.plot(history.history['mse']) # was 'accuracy'])\n        plt.plot(history.history['val_mse']) # was 'val_accuracy'])\n    else :\n        plt.plot(history.history['accuracy'])\n        plt.plot(history.history['val_accuracy'])\n    plt.title('model accuracy')\n    plt.ylabel('accuracy')\n    plt.xlabel('epoch')\n    plt.legend(['train', 'valid'])\n    plt.subplot(222)\n    plt.plot(history.history['loss'])\n    plt.plot(history.history['val_loss'])\n    plt.title('model loss')\n    plt.ylabel('loss')\n    plt.xlabel('epoch')\n    plt.legend(['train', 'valid'])\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:40:44.762021Z","iopub.execute_input":"2023-04-24T17:40:44.762534Z","iopub.status.idle":"2023-04-24T17:40:44.773711Z","shell.execute_reply.started":"2023-04-24T17:40:44.762484Z","shell.execute_reply":"2023-04-24T17:40:44.772579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_his(history, _auto=False)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:40:44.775389Z","iopub.execute_input":"2023-04-24T17:40:44.775833Z","iopub.status.idle":"2023-04-24T17:40:44.827454Z","shell.execute_reply.started":"2023-04-24T17:40:44.775794Z","shell.execute_reply":"2023-04-24T17:40:44.819881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check that Inference Works\n## Test the trained AI with one batch from the Validation set","metadata":{}},{"cell_type":"code","source":"print(\"Let's see how the model behaves on one batch of the validation set\")\nbatch = val_ds.take(1)\none_pred = model.predict(batch)\n\nval_count = (len(batch))\nprint(\"Have classified \", val_count, \"batch of validation images, of batch size \", batch_size)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:41:15.375486Z","iopub.execute_input":"2023-04-24T17:41:15.375956Z","iopub.status.idle":"2023-04-24T17:41:16.55898Z","shell.execute_reply.started":"2023-04-24T17:41:15.375916Z","shell.execute_reply":"2023-04-24T17:41:16.557353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_correct = []\nall_incorrect = []\nfor images, labels in val_ds.take(1):\n    print(\"Here is the shape of one batch of validation examples\\n\", images.shape)\n    image_square = int(np.sqrt(batch_size)) + 1\n    for i in range(batch_size):\n        correct = np.argmax(labels[i])\n        if correct != np.argmax(one_pred[i]):\n            all_incorrect.append(images[i])\n        else:\n            all_correct.append(images[i])\n#         ax = plt.subplot(image_square, image_square, i + 1)\n#         plt.imshow(images[i].numpy().astype(\"uint8\"), cmap='gray', aspect=(num_columns/num_rows), interpolation = 'None')\n#         plt.title(class_names[correct])\n#         plt.axis(\"off\")\n    \n# for i, j in enumerate(all_incorrect) :\n#     print(i, j)\nprint(\"Total images examined \", images.shape[0])\nprint(\"Total images correctly classified \", len(all_correct))\nprint(\"Total images incorrectly classified \", len(all_incorrect))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:41:16.562245Z","iopub.execute_input":"2023-04-24T17:41:16.562759Z","iopub.status.idle":"2023-04-24T17:41:16.741545Z","shell.execute_reply.started":"2023-04-24T17:41:16.56271Z","shell.execute_reply":"2023-04-24T17:41:16.740257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Examine the CNN Filters\n## Visualize each layer of the trained CNN\n* We wish to examine the trained CNN model. In particular, we will be looking to see if we have too many (redundant) convolutional filters, or if there are other issues that can be discerened by looking into the trained CNN.\n* We will do this by traversing the internal workings of the trained CNN weights.\n* First we identify the Conv2D layers, and any MaxPooling layers that precede them\n* NOTE: Here is where I hard-coded the layers, so if you add or remove layers, you'll need to modify the first few lines of the code below","metadata":{}},{"cell_type":"code","source":"###########################################################################################\n# If the user changes the layers of the network, below parameters must also be changed    #\n###########################################################################################\n# NOTE This is manually configured. Need to make this automated.\nnum_conv = 7 # number of convolutional layers\nconv_layer = [2, 3, 5, 7, 9, 10, 12] # the layer (counting from layer 0, the input layer) of the Conv2D's'\nflatten_layer = 13\nembed_layer = 14   # The first Dense layer after the Flatten\nclass_layer = 15   # The last Dense layer (with Softmax Activation) that determines the final classification\nscale_size = [1, 1, 2, 1, 1, 2, 1] # any scaling (due to strides or max pooling) that took place just prior to or at that conv layer (not cumulative)\n###########################################################################################\n# If the user changes the layers of the network, above parameters must also be changed    #\n###########################################################################################\n\n\nmodel_explore = [] # list of models that stop at each convolutional layer\nConv2D_weights = [] # list of filters for convolutional layer\nConv2D_biases = [] # list of filter biases for convolutional layer\n\nfor i in range (num_conv):\n    conv_ind = conv_layer[i]\n    print(\"Conv2D \", i, \"occurs at layer\",conv_ind)\n    model_explore.append(Model(inputs=model.inputs, outputs=model.get_layer(index=conv_ind).output))\n    print(\"Including prior layers, will get an Input Shape of:\",model_explore[i].input_shape)\n    print(\"and have an Output Shape of:\",model_explore[i].output_shape )\n    Conv2D_weights.append(np.copy(model.get_layer(index=conv_ind).get_weights()[0]))\n    Conv2D_biases.append(np.copy(model.get_layer(index=conv_ind).get_weights()[1]))\n    # These are the filters\n    print(\"Final Conv2D has the following rows, columns, color depth, number of filters\", Conv2D_weights[i].shape)\n    # These are the biases\n    print(\"and bias for each filter\", Conv2D_biases[i].shape)\n    print(\"------------------------\")\n\nmodel_flatten = Model(inputs=model.inputs, outputs=model.get_layer(index=flatten_layer).output)\nmodel_embed = Model(inputs=model.inputs, outputs=model.get_layer(index=embed_layer).output)\nmodel_class = Model(inputs=model.inputs, outputs=model.get_layer(index=class_layer).output)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:41:16.743397Z","iopub.execute_input":"2023-04-24T17:41:16.744239Z","iopub.status.idle":"2023-04-24T17:41:16.802306Z","shell.execute_reply.started":"2023-04-24T17:41:16.744186Z","shell.execute_reply":"2023-04-24T17:41:16.801054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create multiple sub-CNN networks\n* Each network will be identical to the original trained network, except the output layer will be one of the Conv2D layers\n* We want to input an image (in this case, features extracted from audio bird calls, transformed into an 8 bit resolution grayscale image) and output the results of that last Conv2D layer.\n* From this intermediate output, we will work backwards through the CNN and attempt to recreate the original image (later in this notebook).\n* We will also use these sub-networks to visualize the CNN filters as described above.","metadata":{}},{"cell_type":"code","source":"# PANv12f now we use a list of models that can be traversed.\nfor conv_l in range(num_conv):\n    print(\"Overall Explore \",conv_l,\" Model Summary:\", model_explore[conv_l].summary(), \"\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:41:16.804352Z","iopub.execute_input":"2023-04-24T17:41:16.805689Z","iopub.status.idle":"2023-04-24T17:41:17.024821Z","shell.execute_reply.started":"2023-04-24T17:41:16.805636Z","shell.execute_reply":"2023-04-24T17:41:17.022625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's observe the internal models of the CNN visually\n* Each convolution filter has weights that form the pattern \"it is looking for.\" \n* More specifically, image areas that most closely match the weights of the convolutional filter will yield a higher numerical value when convolved (multiplied) by that filter.\n* When MaxPooling, it is these most closely matching parts of the image that will have maximal values, and therefore remain, while other, less closely matching portions of the image will be thrown away since they will vield lower convolution values.\n* In this way we can visualize the types of image pixel groupings each CNN filter \"is looking for.\"\n* When multiple CNN layers are connected in sequence, the last layer receives a CNN value from multiple filters in the prior layer. Here too, the last CNN layer is \"looking for\" combinations of prior layer CNN filter convolution results. We can literally combine those prior CNN layer visualizations to form a visualization for each of the last CNN layer filters.\n* In this way, we build up visualizations of what each filter of each CNN layer \"is looking for.\" \n* Layers like Droput are ignored in this analysis\n* Layers like MaxPooling are handled with a scaling factor\n### NOTE: Similar to the \"expansion\" method shown here https://distill.pub/2020/circuits/visualizing-weights/","metadata":{}},{"cell_type":"code","source":"# a list of filter patches, one for each convolutional layer\nnum_filt = []            # number of filters in each Conv2D layer\nfilt_siz = []            # the size of each filter (assume they are square for now)\nborder_siz = []          # the size of the border around the center pixel of the square (assume filters are odd numbered dimensions like 3, 5, 7, etc.)\ncumulative_border = []   # The cumulative effect of the borders from prior Conv3D layers\nfilter_patches = []      # The \"mind reading\" visualization of each filter - equal to the filter at the first Conv2D layer, but more complex later\n\nfor i in range (num_conv):\n    conv_ind = conv_layer[i]\n    print(\"Show the reconstructed image patch for each filter of convolutional layer \", i, \"which occured at layer \",conv_ind, \"in the original model\")\n    num_filt.append(Conv2D_weights[i].shape[3])\n    # filter Size\n    filt_siz.append(Conv2D_weights[i].shape[0])\n    # Border size\n    border_siz.append(int(filt_siz[i] / 2))\n    cb = 0\n    if i == 0 : \n        cumulative_border.append(0)\n    else:\n        for j in range(i) :\n            cb += border_siz[j]\n        cumulative_border.append(cb)\n    print(\"Conv Layer\", i, \" filter Size:\", filt_siz[i], \" x \", filt_siz[i], \n          \" and border size (padding) of \", border_siz[i], \" and cumulative border from prior layers of \", cumulative_border[i])\n    filter_patches.append(np.zeros((filt_siz[i] + 2 * cumulative_border[i], \n                                    filt_siz[i] + 2 * cumulative_border[i], \n                                    num_filt[i])))\n\n    # For now, draw eight filters on a row for display purposes\n    this_num_filt = num_filt[i]\n    if (this_num_filt % 8 == 0) :\n        drows = int(this_num_filt/8)\n    else :\n        drows = int(this_num_filt/8) + 1\n    dcols = 8\n    width = 24\n    height = int(this_num_filt / 4)\n    fig, ax = plt.subplots(nrows=drows, ncols=dcols, figsize=(width, height))\n    \n    thisrow = 0\n    thiscol = 0\n    # temporary variables for filter size and cumulative border size\n    cb = cumulative_border[i]\n    siz = filt_siz[i]\n    image_patch = np.zeros((siz + 2 * cb, siz + 2 * cb))\n    for this_filter in range(this_num_filt) :\n        if i == 0 :\n            image_patch = np.copy(Conv2D_weights[i][:, :, 0, this_filter])\n        else:\n            image_patch.fill(0)\n            for x in range(cb, siz + cb, 1) :\n                for y in range(cb, siz + cb, 1) :\n                    for depth in range(num_filt[i-1]) :\n                        image_patch[x-cb:x+cb+1, y-cb:y+cb+1] += np.copy((Conv2D_weights[i][x-cb, y-cb, depth, this_filter] * \n                                                                         filter_patches[i-1][:,:,depth]))\n    \n        filter_patches[i][:, :, this_filter] = np.copy(np.clip(image_patch, 0, 1e999)) # remove negative numbers (due to \"relu\" activation we used)\n        \n        # display the filter \n        if (drows > 1) :\n            ax[thisrow, thiscol].imshow(filter_patches[i][:, :, this_filter], cmap=\"gray\")\n        else :\n            ax[thiscol].imshow(filter_patches[i][:, :, this_filter], cmap=\"gray\")\n        thiscol += 1\n        if (thiscol >=8) :\n            thisrow += 1\n            thiscol = 0\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:41:17.029957Z","iopub.execute_input":"2023-04-24T17:41:17.030608Z","iopub.status.idle":"2023-04-24T17:41:56.069998Z","shell.execute_reply.started":"2023-04-24T17:41:17.030553Z","shell.execute_reply":"2023-04-24T17:41:56.068418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above shows the filters of each Conv2D layer, multiplied by prior filters, to create a cumilative pattern being targeted by each successive layer (top to bottom)","metadata":{}},{"cell_type":"markdown","source":"## What can we learn from the above?\n* We can see that successive convolution layers manage succesive levels of complexity. These are often described as edge detection first, then these are combined to form textures. Finally, the last convolution layer combines these textures to form patterns it is looking for in order to correctly identify birds. \n* Do we have too many convolutional layers in the middle? Examining the filters, are all the middle layers \"adding value?\"\n* Looking at the last layer (the last group of images) we can see patterns the CNN is looking for. Specifically, look for the white areas and dark areas, and see if the white area is sloped upwards (the bird is going from a high frequency to a low frequency) or downwards (bird goes from low frequency to high frequency, otherwise known as a \"chirp\").\n* Notice also that the specific frequency (musical notes) that the bird is singing is NOT shown at the convolution layer, just the pattern, which might be found at low frequencies or high ones. We would need to reconstruct the signal from the convolutional layer (work backwards through the network) to be able to discern what frequency / time does the pattern occur.\n","metadata":{}},{"cell_type":"markdown","source":"# Observe Data As It Moves Through Each Layer\n## Work backwards through the network to recreate the sounds and original input.\n* We'll just look at the first dozen (12) birds sounds in our batch. \n* For each one, we'll show the original input to the left, and successive recreations of that original input, from left to right - all the way to the last Conv2D layer.\n* Finally, for one of the examples (the last one below) we will take the recreated input (feature vectors) and make a best approximation of the input audio - so the user can hear how much information remains as the NN process the data through each layer.","metadata":{}},{"cell_type":"code","source":"for image in train_ds.take(1) :\n    X_correct = image[0][0:11]\n\n# Now create the output from up to different convolutional layers (using out \"explore\" models)\nprint(\"First examine twelve training examples\")\npred_explore = []\nfor i in range (num_conv):\n    pred_explore.append(model_explore[i].predict(X_correct))\n    print(\"Have processed images through convolutional layer\" ,i, \"of the neural network. Output shape is\", pred_explore[i].shape)\n\n# Let's also look at the outputs of successive Dense layers\npred_flatten = model_flatten.predict(X_correct)\nprint(\"Have processed images through to the Flatten layer of the neural network. Output shape is\", pred_flatten.shape)\n# This is processing the input all the way throug to the first Dense layer after the Flatten (I'll call this the Embedding layer, since it maps all of the input to an n-dimansional vector)\npred_embed = model_embed.predict(X_correct)\nprint(\"Have processed images through to the embedding (first Dense) layer of the neural network. Output shape is\", pred_embed.shape)\n# This is processing the input all the way throug to the last Dense layer (with the Softmax activation) that picks the final classification\npred_class = model_class.predict(X_correct)\nprint(\"Have processed images through to the classification (last Dense with Softmax) layer of the neural network. Output shape is\", pred_class.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:41:56.07166Z","iopub.execute_input":"2023-04-24T17:41:56.072075Z","iopub.status.idle":"2023-04-24T17:41:57.889169Z","shell.execute_reply.started":"2023-04-24T17:41:56.072037Z","shell.execute_reply":"2023-04-24T17:41:57.888233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = X_correct\n# In this section we reverse the convolution layer to recreate a picture\nnum_img = X_test.shape[0]\n\n# display the original and re-created images\nfig, ax = plt.subplots(nrows=num_img, ncols=num_conv+1, figsize=(32, 64))\n\nprint(\"\\n original image on the left followed by successive convolution layer recreations \\n\")\n\nfor i in range(num_img) : # \n    last_scratch = []\n    col = 0\n    # original image\n    ax[i,col].imshow(X_test[i,:,:,0], cmap=\"gray\", aspect=(num_columns/num_rows))\n    last_scratch.append(X_test[i,:,:,:])\n    col += 1\n    for j in range(num_conv) :\n        patch_h = pred_explore[j].shape[1]\n        patch_w = pred_explore[j].shape[2]\n        # Recreate image from combination of this and all of the prior layers of filters, multiplied by the output of this layer\n        tot_border = int((cumulative_border[j] + border_siz[j]))\n        scratch = np.zeros((patch_h + 2 * tot_border,patch_w + 2 * tot_border)) # The size of output plus borders\n        for x in range(tot_border, patch_h+tot_border, 1) :\n            for y in range(tot_border, patch_w+tot_border, 1) :\n                for this_filter in range(num_filt[j]) :\n                    # Find the color patch attributable to this Filter multiplied by the output value it generated\n                    this_color_patch = np.copy(filter_patches[j][:,:,this_filter])\n                    this_color_patch *= (pred_explore[j][i,x-tot_border,y-tot_border,this_filter])               \n                    scratch[x-tot_border:x+tot_border+1, y-tot_border:y+tot_border+1] += np.copy(this_color_patch)\n        PIL_image = Image.fromarray(scratch)\n        scratch = np.array(PIL_image.resize([num_columns, num_rows]))\n        size_h = scratch.shape[0]\n        size_w = scratch.shape[1]\n        last_scratch.append(scratch) # keep a copy of the recreated features for the last image for further analysis\n        ax[i,col].imshow(scratch, cmap=\"gray\", aspect=(num_columns/num_rows))   # here is if we want to show the full reconstruction\n        col += 1\n\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:41:57.89088Z","iopub.execute_input":"2023-04-24T17:41:57.891217Z","iopub.status.idle":"2023-04-24T17:43:18.718594Z","shell.execute_reply.started":"2023-04-24T17:41:57.891185Z","shell.execute_reply":"2023-04-24T17:43:18.71666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above shows 12 different bird sounds (rows) moving through the AI layers (left column is input, remaining columns are successive Conv2D layers)","metadata":{}},{"cell_type":"markdown","source":"## Let's look at the above:\n* We can see that some sort of distinguishing features are carried forwards from the input image (left) through each convolution layer. \n* The last convolution layer (right column) seems the most distorted (fewer visual distinguishing features), and it is indeed the layer before Flatten and Dense.\n* In addition to visual examination of these inner workings of the NN, we can also reverse transform the signal and let a human expert see if the audio still allows correct classification of the bird\n\nLet's do that for the bird in last row above.\n\n## Below, we see and hear each layer of that last row\n* Using that last bird sound (the last row of the above) we will show the recreated input signal, that input signal converted to audio, and both the audio graph and audio playback.\n* We'll look for how much information is retained by the network, in a way that a human expert can tell if sufficient information remains to tell one bird from another.","metadata":{}},{"cell_type":"markdown","source":"# Recreate Input from Data at Every Layer\n## Listen to an approximation of the original input based on output from each layer\n* As you scroll through the below, see and hear how much information remains within each successive layer\n* In particular, can an expert use that information to classify the bird?","metadata":{}},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 0\nprint(\"Original input file.\\n\")\nnorm_trans = np.array(last_scratch[this_layer][:,:,0]) # remove the color dimension (all grayscale)\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:18.720274Z","iopub.execute_input":"2023-04-24T17:43:18.720681Z","iopub.status.idle":"2023-04-24T17:43:19.009638Z","shell.execute_reply.started":"2023-04-24T17:43:18.720642Z","shell.execute_reply":"2023-04-24T17:43:19.008228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:19.012268Z","iopub.execute_input":"2023-04-24T17:43:19.012818Z","iopub.status.idle":"2023-04-24T17:43:22.268723Z","shell.execute_reply.started":"2023-04-24T17:43:19.012769Z","shell.execute_reply":"2023-04-24T17:43:22.267462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:22.270364Z","iopub.execute_input":"2023-04-24T17:43:22.271861Z","iopub.status.idle":"2023-04-24T17:43:22.290245Z","shell.execute_reply.started":"2023-04-24T17:43:22.271806Z","shell.execute_reply":"2023-04-24T17:43:22.288237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 1\nprint(\"Recreated input signal from output of CNN layer, \", this_layer, \" \\n\")\nnorm_trans = np.array(last_scratch[this_layer])\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:22.292152Z","iopub.execute_input":"2023-04-24T17:43:22.292584Z","iopub.status.idle":"2023-04-24T17:43:22.562413Z","shell.execute_reply.started":"2023-04-24T17:43:22.292546Z","shell.execute_reply":"2023-04-24T17:43:22.560953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:22.564459Z","iopub.execute_input":"2023-04-24T17:43:22.565925Z","iopub.status.idle":"2023-04-24T17:43:25.154181Z","shell.execute_reply.started":"2023-04-24T17:43:22.565847Z","shell.execute_reply":"2023-04-24T17:43:25.152729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:25.155574Z","iopub.execute_input":"2023-04-24T17:43:25.155885Z","iopub.status.idle":"2023-04-24T17:43:25.173726Z","shell.execute_reply.started":"2023-04-24T17:43:25.155855Z","shell.execute_reply":"2023-04-24T17:43:25.172148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 2\nprint(\"Recreated input signal from output of CNN layer, \", this_layer, \" \\n\")\nnorm_trans = np.array(last_scratch[this_layer])\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:25.180389Z","iopub.execute_input":"2023-04-24T17:43:25.18077Z","iopub.status.idle":"2023-04-24T17:43:25.448012Z","shell.execute_reply.started":"2023-04-24T17:43:25.180737Z","shell.execute_reply":"2023-04-24T17:43:25.446622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:25.449331Z","iopub.execute_input":"2023-04-24T17:43:25.449693Z","iopub.status.idle":"2023-04-24T17:43:28.292235Z","shell.execute_reply.started":"2023-04-24T17:43:25.449659Z","shell.execute_reply":"2023-04-24T17:43:28.290787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:28.294056Z","iopub.execute_input":"2023-04-24T17:43:28.294396Z","iopub.status.idle":"2023-04-24T17:43:28.311895Z","shell.execute_reply.started":"2023-04-24T17:43:28.294363Z","shell.execute_reply":"2023-04-24T17:43:28.310968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 3\nprint(\"Recreated input signal from output of CNN layer, \", this_layer, \" \\n\")\nnorm_trans = np.array(last_scratch[this_layer])\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:28.313267Z","iopub.execute_input":"2023-04-24T17:43:28.314099Z","iopub.status.idle":"2023-04-24T17:43:28.586531Z","shell.execute_reply.started":"2023-04-24T17:43:28.314064Z","shell.execute_reply":"2023-04-24T17:43:28.585274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:28.588587Z","iopub.execute_input":"2023-04-24T17:43:28.589252Z","iopub.status.idle":"2023-04-24T17:43:31.482281Z","shell.execute_reply.started":"2023-04-24T17:43:28.589215Z","shell.execute_reply":"2023-04-24T17:43:31.480765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:31.484006Z","iopub.execute_input":"2023-04-24T17:43:31.484374Z","iopub.status.idle":"2023-04-24T17:43:31.50232Z","shell.execute_reply.started":"2023-04-24T17:43:31.484339Z","shell.execute_reply":"2023-04-24T17:43:31.501018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 4\nprint(\"Recreated input signal from output of CNN layer, \", this_layer, \" \\n\")\nnorm_trans = np.array(last_scratch[this_layer])\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:31.503972Z","iopub.execute_input":"2023-04-24T17:43:31.504299Z","iopub.status.idle":"2023-04-24T17:43:31.77812Z","shell.execute_reply.started":"2023-04-24T17:43:31.504267Z","shell.execute_reply":"2023-04-24T17:43:31.776783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:31.77982Z","iopub.execute_input":"2023-04-24T17:43:31.780195Z","iopub.status.idle":"2023-04-24T17:43:34.722154Z","shell.execute_reply.started":"2023-04-24T17:43:31.780159Z","shell.execute_reply":"2023-04-24T17:43:34.720699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:34.724011Z","iopub.execute_input":"2023-04-24T17:43:34.724481Z","iopub.status.idle":"2023-04-24T17:43:34.744435Z","shell.execute_reply.started":"2023-04-24T17:43:34.724417Z","shell.execute_reply":"2023-04-24T17:43:34.743132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 5\nprint(\"Recreated input signal from output of CNN layer, \", this_layer, \" \\n\")\nnorm_trans = np.array(last_scratch[this_layer])\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:34.746744Z","iopub.execute_input":"2023-04-24T17:43:34.747651Z","iopub.status.idle":"2023-04-24T17:43:35.023996Z","shell.execute_reply.started":"2023-04-24T17:43:34.747598Z","shell.execute_reply":"2023-04-24T17:43:35.022509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:35.025892Z","iopub.execute_input":"2023-04-24T17:43:35.026356Z","iopub.status.idle":"2023-04-24T17:43:38.405014Z","shell.execute_reply.started":"2023-04-24T17:43:35.026308Z","shell.execute_reply":"2023-04-24T17:43:38.40358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:38.406772Z","iopub.execute_input":"2023-04-24T17:43:38.4079Z","iopub.status.idle":"2023-04-24T17:43:38.425272Z","shell.execute_reply.started":"2023-04-24T17:43:38.40785Z","shell.execute_reply":"2023-04-24T17:43:38.424414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 6\nprint(\"Recreated input signal from output of CNN layer, \", this_layer, \" \\n\")\nnorm_trans = np.array(last_scratch[this_layer])\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:38.426689Z","iopub.execute_input":"2023-04-24T17:43:38.427892Z","iopub.status.idle":"2023-04-24T17:43:38.702855Z","shell.execute_reply.started":"2023-04-24T17:43:38.427854Z","shell.execute_reply":"2023-04-24T17:43:38.701323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:38.704473Z","iopub.execute_input":"2023-04-24T17:43:38.704807Z","iopub.status.idle":"2023-04-24T17:43:41.635307Z","shell.execute_reply.started":"2023-04-24T17:43:38.704775Z","shell.execute_reply":"2023-04-24T17:43:41.633705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:41.637725Z","iopub.execute_input":"2023-04-24T17:43:41.638629Z","iopub.status.idle":"2023-04-24T17:43:41.655647Z","shell.execute_reply.started":"2023-04-24T17:43:41.638571Z","shell.execute_reply":"2023-04-24T17:43:41.654643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the transformed data back into an audio file\nthis_layer = 7\nprint(\"Recreated input signal from output of CNN layer, \", this_layer, \" \\n\")\nnorm_trans = np.array(last_scratch[this_layer])\nprint(\"Shape of image array: \", norm_trans.shape, \", visualization and audio playback\\n\")\nplt.imshow(norm_trans, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:41.657284Z","iopub.execute_input":"2023-04-24T17:43:41.657672Z","iopub.status.idle":"2023-04-24T17:43:41.945938Z","shell.execute_reply.started":"2023-04-24T17:43:41.657637Z","shell.execute_reply":"2023-04-24T17:43:41.944531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", norm_trans.min(), norm_trans.max())\nuntrans_1 = Domain_to_Audio(norm_trans, sr)\nprint(\"This is the recreated audio, (un-transformed) visually displayed as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:41.947794Z","iopub.execute_input":"2023-04-24T17:43:41.948482Z","iopub.status.idle":"2023-04-24T17:43:44.395155Z","shell.execute_reply.started":"2023-04-24T17:43:41.948409Z","shell.execute_reply":"2023-04-24T17:43:44.393769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:44.397112Z","iopub.execute_input":"2023-04-24T17:43:44.397888Z","iopub.status.idle":"2023-04-24T17:43:44.415682Z","shell.execute_reply.started":"2023-04-24T17:43:44.39784Z","shell.execute_reply":"2023-04-24T17:43:44.414381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## NOTE: Best estimate recreations of input can be flawed\nJust because we recreated the original input, doesn't mean our recreation is recognizable by the NN","metadata":{}},{"cell_type":"code","source":"y_this = model.predict(np.reshape(norm_trans, (1, num_rows, num_columns, 1)))\nnorm_trans2 = np.array(last_scratch[0][:,:,0]) # remove the color dimension (all grayscale)\ny_original = model.predict(np.reshape(norm_trans2, (1, num_rows, num_columns, 1)))\n\nprint(\"The above 'recreated input' is not necessarily recognizable by the network as a correct solution.\\n\", \n      \"In this particular case, the reconstructed input above was categorized as bird number \", y_this.argmax(), \"\\n\",\n      \"and the original input was categorized as bird number \", y_original.argmax(), \". \\n\")","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:44.420281Z","iopub.execute_input":"2023-04-24T17:43:44.421497Z","iopub.status.idle":"2023-04-24T17:43:45.20244Z","shell.execute_reply.started":"2023-04-24T17:43:44.421455Z","shell.execute_reply":"2023-04-24T17:43:45.200935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## What can we observe from the above?\n* We can see that with each successive convolutioanl layer, the network is narrowing down to the features that help to distiguish one bird from another - even if it doesn't perfectly the original signal.\n* We can also see the biggest leaps in \"loss\" (when recreating the original signal) take place right after the MaxPool layers\n* Finally, we see that there is a loss of \"time\" information at the last convolutional layer, because after that we flatten the data (row-by-row concatenation) we are blending together the vertical frequency information and the horizontal time information. When we recreate the audio, it sounds like a mish-mash of the sounds the birds make, but not necessarily the cadence / sequence / timing of those sounds.\n\n### Is this a problem?\n* It's really up to you to make that decision. \n* Remember, this is not a human child that has listened to millions of sounds before someone teaches them that \"this is the sound of this type of bird.\" The network was only trained on the sounds of birds, and therefore doesn't know what a human expert knows (such as how to classsify many many other non-bird sounds as well). It may make sense for this \"single use classifier\" to ignore the items mentioned above.\n* Just looking at a blob of \"this bird's sounds\" while ignoring the time sequencing may be an issue if, for example, there are two different birds that use the same musical notes, but in a different pattern broken up by silence over time","metadata":{"execution":{"iopub.status.busy":"2023-04-19T15:19:28.594391Z","iopub.execute_input":"2023-04-19T15:19:28.595052Z","iopub.status.idle":"2023-04-19T15:19:28.604212Z","shell.execute_reply.started":"2023-04-19T15:19:28.594999Z","shell.execute_reply":"2023-04-19T15:19:28.602698Z"}}},{"cell_type":"markdown","source":"# Examine the Final Dense Layers\n* Let's examine what information is provided at the embed layer (the first Dense layer after the Flatten)\n* Let's also examine what information is provided at the classification layer (the last Dense layer with Softmax activation).\n* Here, as before, we are looking at a single bird sound (that we examined in detail above)","metadata":{}},{"cell_type":"code","source":"#######################################################################################################################\n# If changing the shape of the last convolutional later or embed layer, please also change this next lines            #\n# Sorry this is hard coded. Need to make it autoconfigured.                                                           #\n#######################################################################################################################\npre_flatten = (8, 78, 32, 64)   \ninto_flatten = (8, 78, 32)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:43:45.204117Z","iopub.execute_input":"2023-04-24T17:43:45.204496Z","iopub.status.idle":"2023-04-24T17:43:45.210658Z","shell.execute_reply.started":"2023-04-24T17:43:45.204453Z","shell.execute_reply":"2023-04-24T17:43:45.209481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### First see that we can undo the \"Flatten\" layer without loss","metadata":{}},{"cell_type":"code","source":"print(\"First we see that we can undo the Flatten layer\")\nlast_flatten = pred_flatten[num_img - 1]\nprint(\"The input to the Embed layer (the first Dense layer after Flatten) gets the following flattened input shape: \", last_flatten.shape)\nun_flattened = np.reshape(last_flatten, into_flatten)\nprint(\"Un flattened layer shape is \", un_flattened.shape)\n\n# The last convolutional network  is a 3 dimensional array of dimensions rows x columns x filters\nnum_conv_rows = un_flattened.shape[0]\nnum_conv_cols = un_flattened.shape[1]\nnum_conv_filt = un_flattened.shape[2]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer (called embed_data here) multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor j in range(num_conv_filt):\n    this_filter_patch = last_filter_patches[:,:,j]\n    for k in range(border, num_conv_cols + border, 1):\n        for l in range(border, num_conv_rows + border, 1):\n            conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * un_flattened[l-border, k-border, j]\n                \nconv_data2 = np.copy(conv_data1) \nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"\\nIndeed, the flatten layer does not change the best approximation of the input for this bird, given the contents of the last Conv2D layer (flattend and then un-flattened).\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:47:44.285761Z","iopub.execute_input":"2023-04-24T17:47:44.286209Z","iopub.status.idle":"2023-04-24T17:47:44.782934Z","shell.execute_reply.started":"2023-04-24T17:47:44.286174Z","shell.execute_reply":"2023-04-24T17:47:44.781289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now lets see the loss of the Embed layer\nThe embed layer (the first Dense layer after the Flatten) has much fewer neurons (output values) than the prrior Conv2D layer, and so we would expect that data is lost. Let's examine it visually and audibly.\n### Still examining that bird above, where input was passed through each layer","metadata":{}},{"cell_type":"code","source":"last_embed = pred_embed[num_img - 1]\n\n# Creating bar chart of Embed layer output values\nx = np.arange(64)\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, np.array(last_embed))\n\nprint(\"Output of the Embed layer (first Dense layer after Flatten)\")\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:47:44.785513Z","iopub.execute_input":"2023-04-24T17:47:44.786002Z","iopub.status.idle":"2023-04-24T17:47:45.121813Z","shell.execute_reply.started":"2023-04-24T17:47:44.785965Z","shell.execute_reply":"2023-04-24T17:47:45.11998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's calculate the last Conv2D layer for this forced classification\n# Weights from the last Conv2D layer to the flatten->embed layer\nlast_conv_weights_flat = np.copy(model.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the last Conv2D layer to the Embed layer has the shape \", last_conv_weights_flat.shape)\n\nlast_conv_weights = np.copy(np.reshape(last_conv_weights_flat, pre_flatten))\nprint(\"Weights array shape to the embed layer after undoing the Flatten \", last_conv_weights.shape)\n\n# The last convolutional network  is a 3 dimensional arrax of dimensions rows x columns x filters\nnum_conv_rows = last_conv_weights.shape[0]\nnum_conv_cols = last_conv_weights.shape[1]\nnum_conv_filt = last_conv_weights.shape[2]\n# In addition, each of these rows x columns x filter outputs is connected to a dense layer with a number of neurons (called the embed layer)\nnum_embed_neurons = last_conv_weights.shape[3]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor i in range(num_embed_neurons):\n    embed_value = last_embed[i] # using local variables to speed things up\n    for j in range(num_conv_filt):\n        this_filter_patch = last_filter_patches[:,:,j]\n        this_conv_to_embed_weights = last_conv_weights[:,:,j,i]\n        for k in range(border, num_conv_cols + border, 1):\n            for l in range(border, num_conv_rows + border, 1):\n                conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * this_conv_to_embed_weights[l-border,k-border] * embed_value\n                \nfor j in range(num_conv_filt):\n    this_filter_patch = last_filter_patches[:,:,j]\n    for k in range(border, num_conv_cols + border, 1):\n        for l in range(border, num_conv_rows + border, 1):\n            conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * un_flattened[l-border, k-border, j]\n            \n            \nconv_data2 = np.copy(conv_data1) \nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"Here is the best approximation of the input for this bird\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:47:45.123787Z","iopub.execute_input":"2023-04-24T17:47:45.124167Z","iopub.status.idle":"2023-04-24T17:48:00.091208Z","shell.execute_reply.started":"2023-04-24T17:47:45.124133Z","shell.execute_reply":"2023-04-24T17:48:00.089663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", conv_data.min(), conv_data.max())\nuntrans_1 = Domain_to_Audio(conv_data, sr)\nprint(\"This is the recreated audio, (un-transformed) hear and see as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)\nIPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:00.094868Z","iopub.execute_input":"2023-04-24T17:48:00.095335Z","iopub.status.idle":"2023-04-24T17:48:03.242169Z","shell.execute_reply.started":"2023-04-24T17:48:00.095285Z","shell.execute_reply":"2023-04-24T17:48:03.240528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## What can we observe from the above?\n* The first Dense layer (I call it the Embed layer, since it comes just before the Classification layer) seems to be trowing away a great deal of useful information.\n* This is to be expected, because prior to this layer, there were many more numerical values (neuron ouputs) representing the signal, but now this has been reduced down to our embedding dimension.\n* Can you still see and hear distinguishing charactersitics sich as main pitch, rising or lowering pitch over time, etc.?","metadata":{}},{"cell_type":"code","source":"last_class = pred_class[num_img - 1]\n\n# Creating bar chart of Classification layer output values\nx = np.arange(num_birds)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, np.array(last_class))\n\nprint(\"Output of the Classification layer (last Dense layer using Softmax)\")\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:03.244306Z","iopub.execute_input":"2023-04-24T17:48:03.245169Z","iopub.status.idle":"2023-04-24T17:48:03.454975Z","shell.execute_reply.started":"2023-04-24T17:48:03.245114Z","shell.execute_reply":"2023-04-24T17:48:03.453987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# What Does Bird \"X\" Sound Like?\n## What if we don't have an input signal, and just work backwards from a classification?\n* We will repeat this for the first few birds, to see what the \"Robot Mind\" thinks that particular bird sounds like\n","metadata":{}},{"cell_type":"markdown","source":"## Starting with bird class 0 - \"abethr1\"\n* We'll hear what an example original audio sounds like, and see the input features associated with it\n* After that, we'll work backwards through the NN to approximate an input feature set (image) and from that audio best approximation\n* Finally, we'll listen to what the NN has learned each bird \"sounds like\" - you may be surprised!","metadata":{}},{"cell_type":"markdown","source":"### Listen to an example of what that bird sounds like","metadata":{}},{"cell_type":"code","source":"bird_class = 0\nexample_audio_index = 0\naudio_filename = image_data.at[0,\"filename\"]\ncommon_name = image_data.at[0,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n\nprint(\"Here is an example of original audio of this \", common_name, \" bird sounds like\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:03.456321Z","iopub.execute_input":"2023-04-24T17:48:03.457224Z","iopub.status.idle":"2023-04-24T17:48:04.025114Z","shell.execute_reply.started":"2023-04-24T17:48:03.457187Z","shell.execute_reply":"2023-04-24T17:48:04.023722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### For the above example, show the extracted features","metadata":{}},{"cell_type":"code","source":"# These are the features for the above audio sample\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\nprint(\"These are teh features that will be used as input for the above example audio.\\n\")\nplt.imshow(feat, cmap='gray', aspect=int(num_columns/num_rows) , interpolation = 'None')\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:04.027196Z","iopub.execute_input":"2023-04-24T17:48:04.027648Z","iopub.status.idle":"2023-04-24T17:48:04.306136Z","shell.execute_reply.started":"2023-04-24T17:48:04.027609Z","shell.execute_reply":"2023-04-24T17:48:04.305059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Force the output category of the NN to be the above bird","metadata":{}},{"cell_type":"code","source":"# Let's see if we can work backwards from a classification - in this case, we'll use the frst bird\nbird_class = 0 # You can pick a different bird here if you wish\nforce_class = np.zeros(num_birds)\nforce_class[bird_class] = 1\n\nprint(\"Here is the forced output of our classification layer (the last Dense layer with Softmax activation)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_birds)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, force_class)\n\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:04.307701Z","iopub.execute_input":"2023-04-24T17:48:04.308888Z","iopub.status.idle":"2023-04-24T17:48:04.522579Z","shell.execute_reply.started":"2023-04-24T17:48:04.30883Z","shell.execute_reply":"2023-04-24T17:48:04.521002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's calculate the Embed layer for this forced classification\n# Weights from the embed layer to the Classification layer\nclass_weights = np.copy(model_class.get_layer(index=class_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", class_weights.shape)\n\nnum_embed_neurons = class_weights.shape[0]\nnum_class_neurons = class_weights.shape[1]\nembed_data = np.zeros(num_embed_neurons)\nfor i in range(num_embed_neurons):\n    for j in range(num_class_neurons):\n        embed_data[i] += force_class[j] * class_weights[i, j]\n\nprint(\"Here is the forced output of our embed layer (the first Dense layer after flattened last Conv layer)\")\n\n# Creating bar chart of Embed layer output values\nx = np.arange(num_embed_neurons)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, embed_data)\n\n# Show plot\nplt.show()\n\n# Weights from the embed layer to the Flatten-ed last Conv layer\nembed_weights = np.copy(model_embed.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the flattened last Conv layer to the embed has the shape \", embed_weights.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:04.524535Z","iopub.execute_input":"2023-04-24T17:48:04.525036Z","iopub.status.idle":"2023-04-24T17:48:04.837852Z","shell.execute_reply.started":"2023-04-24T17:48:04.524987Z","shell.execute_reply":"2023-04-24T17:48:04.836793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Work backwards through the NN to recreate a best approximation of the input features, and audio, of the above bird.\n* Essentially \"What has the NN learned to be the audio feature set for the above bird, and what does it sound like?\"","metadata":{}},{"cell_type":"code","source":"# Let's calculate the last Conv2D layer for this forced classification\n# Weights from the last Conv2D layer to the embed layer\nlast_conv_weights_flat = np.copy(model.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", last_conv_weights_flat.shape)\n\nlast_conv_weights = np.copy(np.reshape(last_conv_weights_flat, pre_flatten))\nprint(\"Weights array shape from the embed layer after undoing the Flatten \", last_conv_weights.shape)\n\n# The last convolutional network  is a 3 dimensional arrax of dimensions rows x columns x filters\nnum_conv_rows = last_conv_weights.shape[0]\nnum_conv_cols = last_conv_weights.shape[1]\nnum_conv_filt = last_conv_weights.shape[2]\n# In addition, each of these rows x columns x filter outputs is connected to a dense layer with a number of neurons (called the embed layer)\nnum_embed_neurons = last_conv_weights.shape[3]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer (called embed_data here) multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor i in range(num_embed_neurons):\n    embed_value = embed_data[i] # using local variables to speed things up\n    for j in range(num_conv_filt):\n        this_filter_patch = last_filter_patches[:,:,j]\n        this_conv_to_embed_weights = last_conv_weights[:,:,j,i]\n        for k in range(border, num_conv_cols + border, 1):\n            for l in range(border, num_conv_rows + border, 1):\n                conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * this_conv_to_embed_weights[l-border,k-border] * embed_value\n                \n\nconv_data2 = np.copy(conv_data1)\nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"Here is the best approximation of the input for this bird\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:04.8447Z","iopub.execute_input":"2023-04-24T17:48:04.845184Z","iopub.status.idle":"2023-04-24T17:48:19.228097Z","shell.execute_reply.started":"2023-04-24T17:48:04.845142Z","shell.execute_reply":"2023-04-24T17:48:19.227216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = model.predict(np.reshape(conv_data, (1, num_rows, num_columns, 1)))\n\nprint(\"As noted above, recreation of input does not mean that recreation will be classified the same way. \\n\",\n      \"In fact, using the above as input, the NN categorizes this as bird number \", x.argmax())","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:19.229368Z","iopub.execute_input":"2023-04-24T17:48:19.22992Z","iopub.status.idle":"2023-04-24T17:48:19.309469Z","shell.execute_reply.started":"2023-04-24T17:48:19.229886Z","shell.execute_reply":"2023-04-24T17:48:19.308144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", conv_data.min(), conv_data.max())\nuntrans_1 = Domain_to_Audio(conv_data, sr)\nprint(\"This is the recreated audio, (un-transformed) hear and see as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)\nIPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:19.311093Z","iopub.execute_input":"2023-04-24T17:48:19.313785Z","iopub.status.idle":"2023-04-24T17:48:21.800709Z","shell.execute_reply.started":"2023-04-24T17:48:19.313733Z","shell.execute_reply":"2023-04-24T17:48:21.799112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above is what the AI \"thinks bird class 0 - abethr1 - sounds like.\" \nIf you'll forgive the personification, pretending like the AI is actually thinking\n### What can we learn from the above?\n* We can see that certain important frequencies appear in the same position (in the extracted feature image)\n* We can see that certain important cadences (low to high notes, versus high to low notes) can be spotted\n* Some information, such as rhythm/timing of repeated sounds, seems to be completely lost.\n* We see that quite a bit of information has been thrown away (I can't tell what bird it is from the audio, and probabbly neither can an expert).\n* Could this discarded information be important? Experts would say \"yes\" but AI researchers migth be OK with a good score only.","metadata":{}},{"cell_type":"markdown","source":"## Next is bird class 1 - \"abhori1\"\n* We'll hear what an example original audio sounds like\n* After that, we'll work backwards through the NN to approximate an input feature set (image) and from that audio best approximation\n* Listen to what the NN has learned each bird \"sounds like\" - you may be surprised!","metadata":{}},{"cell_type":"markdown","source":"### Listen to an example of what that bird sounds like","metadata":{}},{"cell_type":"code","source":"bird_class = 1\nexample_audio_index = 15\naudio_filename = image_data.at[example_audio_index,\"filename\"]\ncommon_name = image_data.at[example_audio_index,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n\nprint(\"Here is an example of original audio of this \", common_name, \" bird sounds like\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:21.802228Z","iopub.execute_input":"2023-04-24T17:48:21.802587Z","iopub.status.idle":"2023-04-24T17:48:22.347514Z","shell.execute_reply.started":"2023-04-24T17:48:21.802553Z","shell.execute_reply":"2023-04-24T17:48:22.346621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### For the above example, show the extracted features","metadata":{}},{"cell_type":"code","source":"# These are the features for the above audio sample\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\nprint(\"These are the features that will be used as input for the above example audio.\\n\")\nplt.imshow(feat, cmap='gray', aspect=int(num_columns/num_rows) , interpolation = 'None')\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:22.348664Z","iopub.execute_input":"2023-04-24T17:48:22.349458Z","iopub.status.idle":"2023-04-24T17:48:22.611317Z","shell.execute_reply.started":"2023-04-24T17:48:22.349409Z","shell.execute_reply":"2023-04-24T17:48:22.609595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Force the output category of the NN to be the above bird","metadata":{}},{"cell_type":"code","source":"# Let's see if we can work backwards from a classification - in this case, we'll use the frst bird\nbird_class = 1 # You can pick a different bird here if you wish\nforce_class = np.zeros(num_birds)\nforce_class[bird_class] = 1\n\nprint(\"Here is the forced output of our classification layer (the last Dense layer with Softmax activation)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_birds)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, force_class)\n\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:22.613095Z","iopub.execute_input":"2023-04-24T17:48:22.613485Z","iopub.status.idle":"2023-04-24T17:48:22.813599Z","shell.execute_reply.started":"2023-04-24T17:48:22.613416Z","shell.execute_reply":"2023-04-24T17:48:22.812399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's calculate the Embed layer for this forced classification\n# Weights from the embed layer to the Classification layer\nclass_weights = np.copy(model_class.get_layer(index=class_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", class_weights.shape)\n\nnum_embed_neurons = class_weights.shape[0]\nnum_class_neurons = class_weights.shape[1]\nembed_data = np.zeros(num_embed_neurons)\nfor i in range(num_embed_neurons):\n    for j in range(num_class_neurons):\n        embed_data[i] += force_class[j] * class_weights[i, j]\n\nprint(\"Here is the forced output of our embed layer (the first Dense layer after flattened last Conv layer)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_embed_neurons)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, embed_data)\n\n# Show plot\nplt.show()\n\n# Weights from the embed layer to the Flatten-ed last Conv layer\nembed_weights = np.copy(model_embed.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the flattened last Conv layer to the embed has the shape \", embed_weights.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:22.81493Z","iopub.execute_input":"2023-04-24T17:48:22.815251Z","iopub.status.idle":"2023-04-24T17:48:23.146001Z","shell.execute_reply.started":"2023-04-24T17:48:22.815219Z","shell.execute_reply":"2023-04-24T17:48:23.144477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Work backwards through the NN to recreate a best approximation of the input features, and audio, of the above bird.\n* Essentially \"What has the NN learned to be the audio feature set for the above bird, and what does it sound like?\"","metadata":{}},{"cell_type":"code","source":"# Let's calculate the last Conv2D layer for this forced classification\n# Weights from the last Conv2D layer to the embed layer\nlast_conv_weights_flat = np.copy(model.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", last_conv_weights_flat.shape)\n\nlast_conv_weights = np.copy(np.reshape(last_conv_weights_flat, pre_flatten))\nprint(\"Weights array shape from the embed layer after undoing the Flatten \", last_conv_weights.shape)\n\n# The last convolutional network  is a 3 dimensional arrax of dimensions rows x columns x filters\nnum_conv_rows = last_conv_weights.shape[0]\nnum_conv_cols = last_conv_weights.shape[1]\nnum_conv_filt = last_conv_weights.shape[2]\n# In addition, each of these rows x columns x filter outputs is connected to a dense layer with a number of neurons (called the embed layer)\nnum_embed_neurons = last_conv_weights.shape[3]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer (called embed_data here) multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor i in range(num_embed_neurons):\n    embed_value = embed_data[i] # using local variables to speed things up\n    for j in range(num_conv_filt):\n        this_filter_patch = last_filter_patches[:,:,j]\n        this_conv_to_embed_weights = last_conv_weights[:,:,j,i]\n        for k in range(border, num_conv_cols + border, 1):\n            for l in range(border, num_conv_rows + border, 1):\n                conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * this_conv_to_embed_weights[l-border,k-border] * embed_value\n                \nconv_data2 = np.copy(conv_data1) # [border:num_conv_rows + border, border:num_conv_cols + border])\nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"Here is the best approximation of the input for this bird\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:23.148197Z","iopub.execute_input":"2023-04-24T17:48:23.149243Z","iopub.status.idle":"2023-04-24T17:48:37.419733Z","shell.execute_reply.started":"2023-04-24T17:48:23.149192Z","shell.execute_reply":"2023-04-24T17:48:37.418234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = model.predict(np.reshape(conv_data, (1, num_rows, num_columns, 1)))\n\nprint(\"As noted above, recreation of input does not mean that recreation will be classified the same way. \\n\",\n      \"Using the above as input, the NN categorizes this as bird number \", x.argmax())","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:37.421574Z","iopub.execute_input":"2023-04-24T17:48:37.42366Z","iopub.status.idle":"2023-04-24T17:48:37.508692Z","shell.execute_reply.started":"2023-04-24T17:48:37.423602Z","shell.execute_reply":"2023-04-24T17:48:37.507308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", conv_data.min(), conv_data.max())\nuntrans_1 = Domain_to_Audio(conv_data, sr)\nprint(\"This is the recreated audio, (un-transformed) hear and see as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)\nIPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:37.511816Z","iopub.execute_input":"2023-04-24T17:48:37.512246Z","iopub.status.idle":"2023-04-24T17:48:40.15522Z","shell.execute_reply.started":"2023-04-24T17:48:37.51221Z","shell.execute_reply":"2023-04-24T17:48:40.153952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above is what the AI \"thinks bird class 1 - abhori1 - sounds like.\"\n* Many of the same observations regarding retained vversus discarded information","metadata":{}},{"cell_type":"markdown","source":"## Finally bird class 2 - \"abythr1\"\n* We'll hear what an example original audio sounds like\n* After that, we'll work backwards through the NN to approximate an input feature set (image) and from that audio best approximation\n* Listen to what the NN has learned each bird \"sounds like\" - you may be surprised!","metadata":{}},{"cell_type":"markdown","source":"### Listen to an example of what that bird sounds like","metadata":{}},{"cell_type":"code","source":"bird_class = 2\nexample_audio_index = 141\naudio_filename = image_data.at[example_audio_index,\"filename\"]\ncommon_name = image_data.at[example_audio_index,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n\nprint(\"Here is an example of original audio of this \", common_name, \" bird sounds like\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:40.156842Z","iopub.execute_input":"2023-04-24T17:48:40.157171Z","iopub.status.idle":"2023-04-24T17:48:40.703566Z","shell.execute_reply.started":"2023-04-24T17:48:40.157138Z","shell.execute_reply":"2023-04-24T17:48:40.702125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### For the above example, show the extracted features","metadata":{}},{"cell_type":"code","source":"# These are the features for the above audio sample\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\nprint(\"These are teh features that will be used as input for the above example audio.\\n\")\nplt.imshow(feat, cmap='gray', aspect=int(num_columns/num_rows) , interpolation = 'None')\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:40.705821Z","iopub.execute_input":"2023-04-24T17:48:40.706221Z","iopub.status.idle":"2023-04-24T17:48:41.449851Z","shell.execute_reply.started":"2023-04-24T17:48:40.706184Z","shell.execute_reply":"2023-04-24T17:48:41.448357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Force the output category of the NN to be the above bird","metadata":{}},{"cell_type":"code","source":"# Let's see if we can work backwards from a classification - in this case, we'll use the frst bird\nbird_class = 2 # You can pick a different bird here if you wish\nforce_class = np.zeros(num_birds)\nforce_class[bird_class] = 1\n\nprint(\"Here is the forced output of our classification layer (the last Dense layer with Softmax activation)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_birds)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, force_class)\n\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:41.451536Z","iopub.execute_input":"2023-04-24T17:48:41.451902Z","iopub.status.idle":"2023-04-24T17:48:41.662801Z","shell.execute_reply.started":"2023-04-24T17:48:41.451865Z","shell.execute_reply":"2023-04-24T17:48:41.661027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's calculate the Embed layer for this forced classification\n# Weights from the embed layer to the Classification layer\nclass_weights = np.copy(model_class.get_layer(index=class_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", class_weights.shape)\n\nnum_embed_neurons = class_weights.shape[0]\nnum_class_neurons = class_weights.shape[1]\nembed_data = np.zeros(num_embed_neurons)\nfor i in range(num_embed_neurons):\n    for j in range(num_class_neurons):\n        embed_data[i] += force_class[j] * class_weights[i, j]\n\nprint(\"Here is the forced output of our embed layer (the first Dense layer after flattened last Conv layer)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_embed_neurons)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, embed_data)\n\n# Show plot\nplt.show()\n\n# Weights from the embed layer to the Flatten-ed last Conv layer\nembed_weights = np.copy(model_embed.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the flattened last Conv layer to the embed has the shape \", embed_weights.shape)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:41.664451Z","iopub.execute_input":"2023-04-24T17:48:41.664807Z","iopub.status.idle":"2023-04-24T17:48:41.972861Z","shell.execute_reply.started":"2023-04-24T17:48:41.664773Z","shell.execute_reply":"2023-04-24T17:48:41.971443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Work backwards through the NN to recreate a best approximation of the input features, and audio, of the above bird.\n* Essentially \"What has the NN learned to be the audio feature set for the above bird, and what does it sound like?\"","metadata":{}},{"cell_type":"code","source":"# Let's calculate the last Conv2D layer for this forced classification\n# Weights from the last Conv2D layer to the embed layer\nlast_conv_weights_flat = np.copy(model.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", last_conv_weights_flat.shape)\n\nlast_conv_weights = np.copy(np.reshape(last_conv_weights_flat, pre_flatten))\nprint(\"Weights array shape from the embed layer after undoing the Flatten \", last_conv_weights.shape)\n\n# The last convolutional network  is a 3 dimensional arrax of dimensions rows x columns x filters\nnum_conv_rows = last_conv_weights.shape[0]\nnum_conv_cols = last_conv_weights.shape[1]\nnum_conv_filt = last_conv_weights.shape[2]\n# In addition, each of these rows x columns x filter outputs is connected to a dense layer with a number of neurons (called the embed layer)\nnum_embed_neurons = last_conv_weights.shape[3]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer (called embed_data here) multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation.\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor i in range(num_embed_neurons):\n    embed_value = embed_data[i] # using local variables to speed things up\n    for j in range(num_conv_filt):\n        this_filter_patch = last_filter_patches[:,:,j]\n        this_conv_to_embed_weights = last_conv_weights[:,:,j,i]\n        for k in range(border, num_conv_cols + border, 1):\n            for l in range(border, num_conv_rows + border, 1):\n                conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * this_conv_to_embed_weights[l-border,k-border] * embed_value\n                \nconv_data2 = np.copy(conv_data1)\nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"Here is the best approximation of the input for this bird\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:41.974772Z","iopub.execute_input":"2023-04-24T17:48:41.975096Z","iopub.status.idle":"2023-04-24T17:48:56.35353Z","shell.execute_reply.started":"2023-04-24T17:48:41.975065Z","shell.execute_reply":"2023-04-24T17:48:56.352043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = model.predict(np.reshape(conv_data, (1, num_rows, num_columns, 1)))\n\nprint(\"As noted above, recreation of input does not mean that recreation will be classified the same way. \\n\",\n      \"In fact, using the above as input, the NN categorizes this as bird number \", x.argmax())\n","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:56.355176Z","iopub.execute_input":"2023-04-24T17:48:56.355572Z","iopub.status.idle":"2023-04-24T17:48:56.444124Z","shell.execute_reply.started":"2023-04-24T17:48:56.355536Z","shell.execute_reply":"2023-04-24T17:48:56.442541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", conv_data.min(), conv_data.max())\nuntrans_1 = Domain_to_Audio(conv_data, sr)\nprint(\"This is the recreated audio, (un-transformed) hear and see as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)\nIPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:56.44622Z","iopub.execute_input":"2023-04-24T17:48:56.446612Z","iopub.status.idle":"2023-04-24T17:48:59.500616Z","shell.execute_reply.started":"2023-04-24T17:48:56.446579Z","shell.execute_reply":"2023-04-24T17:48:59.499568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above is what the AI \"thinks bird class 2 - abythr1 - sounds like.\"\n* Many of the same observations regarding retained vversus discarded information","metadata":{}},{"cell_type":"markdown","source":"# Submit the Results\n* You will likely run this notebook twice or three times - first to create all the training set images, then to train a NN, and finally to use that trained NN to infer solutions to the contest.\n* I have included all of the code - including submission to the contest - to make this easier for you.\n\nBelow code modified from \"Inferring Birds with Kaggle Models\", courtesy PHIL CULLITON +1 · COPIED FROM PRIVATE NOTEBOOK +42,-5 · 1MO AGO · 22,483 VIEWS","metadata":{}},{"cell_type":"code","source":"def predict_for_sample(filename, sample_submission, frame_limit_secs=None):\n    file_id = filename.split(\".ogg\")[0].split(\"/\")[-1]\n    \n    audio, sample_rate = librosa.load(filename, sr = sr)\n        \n    # Get number of samples for 5 seconds\n    buffer = segment\n\n    samples_total = len(audio)\n    samples_segmented = 0\n    counter = 1\n\n    frame = 5\n\n    # this will be a list of 5 second segments\n    feature_split = []\n    classifications = []\n\n    if num_birds < len(competition_class_map):\n        print(\"Have to append 0% likelihood bird types because only \", num_birds, \"were used in taining\\n\")\n\n    while samples_segmented < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_segmented):\n            buffer = samples_total - samples_segmented\n\n        block = audio[samples_segmented : (samples_segmented + buffer)]\n\n        # convert the features \n        these_features = Audio_to_Domain(block, sr)\n\n        #scale to 0 - 255\n        fmax = these_features.max()\n        fmin = these_features.min()\n        frange = fmax - fmin\n            \n        these_features = np.array((((these_features - fmin) / frange)*255))\n        feature_split.append(these_features)\n\n        # Classfy each 5 secons segment\n        scratch = np.copy(np.reshape(these_features, (1,num_rows, num_columns, 1)))\n        scratch2 = (model.predict(scratch, verbose=0))\n        classifications.append(scratch2)\n        \n        counter += 1\n        samples_segmented += buffer\n                \n        probabilities = tf.nn.softmax(scratch2).numpy()\n        probabilities = np.copy(np.reshape(probabilities, (num_birds)))\n        if num_birds < len(competition_class_map) :\n            probabilities = np.append(probabilities, np.zeros(len(competition_class_map) - num_birds))\n            print(\"Had to append bird types because only \", num_birds, \"were used in taining\")\n        print(samples_segmented, samples_total)\n        \n        ## set the appropriate row in the sample submission\n        sample_submission.loc[sample_submission.row_id == file_id + \"_\" + str(frame), competition_classes] = probabilities[competition_class_map]\n        frame += 5\n        \ndef predict_for_sample(filename, sample_submission, frame_limit_secs=None):\n    file_id = filename.split(\".ogg\")[0].split(\"/\")[-1]\n    \n    audio, sample_rate = librosa.load(filename, sr = sr)\n        \n    # Get number of samples for 5 seconds\n    buffer = segment\n\n    samples_total = len(audio)\n    samples_segmented = 0\n    counter = 1\n\n    frame = 5\n\n    # this will be a list of 5 second segments\n    feature_split = []\n    classifications = []\n\n    if num_birds < len(competition_class_map):\n        print(\"Had to append bird types because only \", num_birds, \"were used in taining\")\n    while samples_segmented < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_segmented):\n            buffer = samples_total - samples_segmented\n\n        block = audio[samples_segmented : (samples_segmented + buffer)]\n\n        # convert the features \n        these_features = Audio_to_Domain(block, sr)\n\n        #scale to 0 - 255\n        fmax = these_features.max()\n        fmin = these_features.min()\n        frange = fmax - fmin\n            \n        these_features = np.array((((these_features - fmin) / frange)*255))\n        feature_split.append(these_features)\n\n        # Classfy each 5 second segment\n        scratch = np.copy(np.reshape(these_features, (1, num_rows, num_columns, 1)))\n        scratch2 = (model.predict(scratch, verbose = 1))\n        classifications.append(scratch2)\n        \n        counter += 1\n        samples_segmented += buffer\n                \n        probabilities = tf.nn.softmax(scratch2).numpy()\n        probabilities = np.copy(np.reshape(probabilities, (num_birds)))\n        if num_birds < len(competition_class_map) :\n            probabilities = np.append(probabilities, np.zeros(len(competition_class_map) - num_birds))\n        # print(samples_segmented, samples_total)\n        \n        ## set the appropriate row in the sample submission\n        sample_submission.loc[sample_submission.row_id == file_id + \"_\" + str(frame), competition_classes] = probabilities[competition_class_map]\n        frame += 5","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:59.502316Z","iopub.execute_input":"2023-04-24T17:48:59.503043Z","iopub.status.idle":"2023-04-24T17:48:59.525331Z","shell.execute_reply.started":"2023-04-24T17:48:59.503005Z","shell.execute_reply":"2023-04-24T17:48:59.523988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\ntest_samples = list(glob.glob(\"/kaggle/input/birdclef-2023/test_soundscapes/*.ogg\"))\ntest_samples","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:59.532246Z","iopub.execute_input":"2023-04-24T17:48:59.532705Z","iopub.status.idle":"2023-04-24T17:48:59.549567Z","shell.execute_reply.started":"2023-04-24T17:48:59.532666Z","shell.execute_reply":"2023-04-24T17:48:59.547882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_metadata = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\ncompetition_classes = sorted(train_metadata.primary_label.unique())\n\ncompetition_class_map = []\nfor c in competition_classes:\n        i = competition_classes.index(c)\n        competition_class_map.append(i)\n        \n\nsample_sub = pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\nsample_sub[competition_classes] = sample_sub[competition_classes].astype(np.float32)\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:59.551615Z","iopub.execute_input":"2023-04-24T17:48:59.552132Z","iopub.status.idle":"2023-04-24T17:48:59.73458Z","shell.execute_reply.started":"2023-04-24T17:48:59.552082Z","shell.execute_reply":"2023-04-24T17:48:59.733032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frame_limit_secs = 15 if sample_sub.shape[0] == 3 else None\nfor sample_filename in test_samples:\n    predict_for_sample(sample_filename, sample_sub, frame_limit_secs=15)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:48:59.738655Z","iopub.execute_input":"2023-04-24T17:48:59.739731Z","iopub.status.idle":"2023-04-24T17:49:24.876501Z","shell.execute_reply.started":"2023-04-24T17:48:59.739691Z","shell.execute_reply":"2023-04-24T17:49:24.875148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:49:24.878469Z","iopub.execute_input":"2023-04-24T17:49:24.87924Z","iopub.status.idle":"2023-04-24T17:49:24.915018Z","shell.execute_reply.started":"2023-04-24T17:49:24.879192Z","shell.execute_reply":"2023-04-24T17:49:24.913103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-04-24T17:49:24.917138Z","iopub.execute_input":"2023-04-24T17:49:24.917612Z","iopub.status.idle":"2023-04-24T17:49:24.936495Z","shell.execute_reply.started":"2023-04-24T17:49:24.917575Z","shell.execute_reply":"2023-04-24T17:49:24.935093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Additional Background Information\nAs you know, in this notebook, we are taking the first step to generate a submission to the [BirdClef2023 competition](https://www.kaggle.com/c/birdclef-2023).  The goal of the competition is to identify Eastern African bird species by sound. \n### \"Reading the Robot Mind\" is a system that allows non-programmers to examine performance of individual and combined stages of the classification process\nMost important for this system is that it can be accomplished by a subject matter expert who may not be familair with the inner workings of AI. Subject matter experts know how to listen to birds and identify them, so they need data presented as audio (and possibly spectral analysis, as I've seen this used extensively by experts when exmining data). Additional references and history of the reading the robot mind system is explained in detail below. \n#### Stages of an automated process to classify birds:\nThe stages of the desired classification system include:\n\n* A method to record audio (provided with the contest)\n* A method of segmenting audio data (somewhat provided with the contest - 5 second segments)\n* A method of transforming the audio segment into a sequence of feature vectors (feature extraction such as Mel Scale Frequency Spectrum and MFCC)\n* A method of converting sequence of features into an image for CNN processing (in this case, also quantizing to 8-bit unsigned values)\n* A method where several CNN (Conv2D) and data reduction (MaxPooling) layers process data in sequence, iteratively extracting salient aspects\n* A method of taking a multi-layer CNN and outputting an embedding (Flatten then Dense layer)\n* A method of taking the embedding and converting it to a bird classification (Dense with softmax activation)\n\nPrior to the \"Reading the Robot Mind\" system described herein; a subject matter expert will not be able to examine issues associated with individual stages (such as listening to audio and looking at visualizations) except for the first two stages. Issues that can arise include: Inability to observe the difference between samples from different bird classifications. Inability to observe similarities between samples from the same bird classification. Inability to use audio and visual clues to correctly classify the bird.\n\nThis notebook presents a method whereby at any stage of the classification system, the expert can examine the internal workings of the system, and work backwards to audio and visual data in a way that is interpretable by a human expert.\n### We will qualitatively examine feature extraction algorithms as part of an effort to \"Read the Robot Mind\"\nThis notebook creates a method to allow users to select different hyper-paramenters (methods) at each stage and qualitatively compare them. Variations of segmentation algorithms are not presented in this notebook (since they are dictated to some extent by the rules of this competition).\n### Also examined is the Neural Network we trained. Can it too be used to \"Read the Robot Mind?\"\nHere we train a simple NN and see that it can classify with some level of accuracy. We can then use that NN \"in reverse\" to attempt to recreate an audio signal that generated a particular stage (often referred to as a layer) in the classification.\n### Finally - what if we simply supply a classification?\nThis notebook demonstrates the ability to simply provide the output - and then work backwards through the AI layers and reproduce a best approximation of the input that cause this classification. Effectively asking the AI system to \"Let me hear what you've learned a particular bird sounds like.\"","metadata":{}},{"cell_type":"markdown","source":"# Brief History of the \"Reading the Robot Mind\" System\nIt used to be that only human experts examined data and made decisions. Now Artificial Intelligence (AI) is enabling robotic decision making in an ever-widening variety of applications. As society allows this to happen, there is a greater likelihood that these robot decisions can affect people’s lives. It makes sense, therefore, to understand the capabilities and societal implications of AI robots.\n\nBig data is a term used to describe both the opportunities and the problems associated with so much information now available for decision making. With the advent of the Internet of Things (IoT) the impact of this huge amount of data is only growing. Actionable decisions need to be distilled from big data and AI can only go so far based on statistical analysis. Because of this, many deep learning algorithms are being developed. \n\n##  What exactly have these AI robots learned so deeply from all of this big data? \n\nThis question is very reasonable for society to ask. It is not enough to train and create a great AI robot. Many researchers are realizing that before their systems can be deployed, they must be able to prove to human experts that the robots learned the right things from the right data. This is difficult because human expert decision makers are not necessarily the same people who are good at creating robot AI. These two teams must work together in a user friendly way.\n## If only we could read the robot mind.\nRecently some old research of mine has been getting increased attention by engineers and scientists working on exactly the issues discussed above. This research yielded four US patents and a Windows app that gives a friendly user interface to allow examination of big data and inner workings of a trained AI decision maker. I never publically released the software, called INTEGRAT, because I sold rights to the patents to my employer at the time. \n\n## Now that the patents have expired, I can share my work.\n\n* For even more details, the reader is directed to the specific patent and independent claims listed below, all publically available at www.USPTO.gov.\n","metadata":{}},{"cell_type":"markdown","source":"# US Patent 5809462\nPaul A. Nussbaum, Inventor\n## Question: Is the feature extraction algorithm good and can the features be used as a codec?\n(\"codec\" is a portmanteau or combination of the words \"coder\" and \"decoder\", such as jpeg or mp3, etc...)\n\nAll AI systems train on data from the real world gathered by sensors, but before the data can make it to the AI, it must be digitized and important features must be extracted (sometimes called data transformation). \"Reading the Robot Mind\" allows a human expert to determine if the features extracted have retained sufficient information to identify the correct classification. \n\n* The question is: does the feature extraction algorithm retain classification-necessary information (or have we thrown away too much information)?\n\n* The premise is as follows: Human experts can observe the original signal and/or measurements and visualizations thereof - and correctly make the classification. \n\n* The thesis is that: A sufficient condition that the feature extraction algorithm isn't throwing away too much info as it seeks to condition and compress the signal; is that there exists a reverse transformation able to use the extracted features to recreate the original signal and/or measurements and visualizations thereof - such that the human expert will use these to correctly make the classification.\n\n* A bonus is: The human expert can also be presented the extracted features, and may find that they can be useful in future human classification exercises (observedly similar for same classifications, while different from others).\n\nSalient points from US Patent 5809462:\n\n* Independent Claim 1 – Provide a friendly user interface to present the features to the eyes and ears of the human expert to see if they can make the correct decisions based on the transformation features alone.\n   \n* Independent Claim 14 – Use that same user interface to let the human expert make sure that the data used to train one possible decision all seem similar to each other, and different from data which should yield different decisions.\n   \n* Independent Claim 25 – Let the human expert see if the features contain enough information to be used as a codec. In other words, if we can transform the original data into a set of features, we should be able to create original data (somewhat distorted and simplified perhaps) from the features – with sufficient information for a human expert to use that to identify the correct decision. \n","metadata":{}},{"cell_type":"markdown","source":"# US Patent 5867816\nPaul A. Nussbaum, Inventor\n## Question: Will the robot function just as well when the data segmentation, feature extraction, and decision identifier are fully automated?\n\nThe AI Robot is not useful if it requires a human expert to accompany the device when it is deployed in the field for day to day use. The robot must work on its own. Nevertheless, a quality assurance mechanism is required that makes the robot demonstrate it is working in “fully automatic” mode while a human expert grades it on the decisions the AI has made.\n\n* \"Reading the Robot Mind\" means that this segmentation, feature extraction, and classifier pipeline (inference) are observable in situ (on-site in real deployed situations, preferably live).\n\n* Bonus: The human expert can make adjustments to this pipeline to help get the desired classification accuracy, without requiring technical knowledge of the pipeline.\n\nIn the BirdClef2023 \n\nThe salient point from US Patent 5867816 is:\n\n* Independent Claims 1, 11, and 29 – The human expert examines and modifies the segmentation, features, and decision identifications to improve automated functioning.\n","metadata":{}},{"cell_type":"markdown","source":"# US Patent 5749066\nPaul A. Nussbaum, Inventor\n## Question: Is the training data good and can the classifier be used as a codec? \n\nAI robots can be finicky about the actual examples selected to be used for training. Giving too many of one kind of example can have a negative impact on the resulting decision maker, as can the introduction of misleading data examples (either accidentally or maliciously). Sometimes, a training set can be obviously bad, and other times, only after the AI is trained, can the human expert see that something has gone wrong.\n   \n* Independent Claims 1 and 7 - Provide a friendly user interface to present all stages of creating the training set to the eyes and ears of the human expert to see if they can spot any problems noted above.\n\n* Independent Claim 3 – Let the human expert see if the AI can be used as a codec. In other words, if the AI can identify the correct decision, then we can work backwards through the AI layers to recreate an exemplary feature set. We can then create the original data (somewhat distorted and simplified perhaps) from those features – with sufficient information for a human expert to use that to identify the correct decision.\n","metadata":{}},{"cell_type":"markdown","source":"#  US Patent 5864803\nThis Notebook DOES NOT cover these aspects in detail\n## Question: How can we assign more than one correct classification when experts cannot agree?\nBecause data sets can be very large and continuously streaming in from many sources, AI robot creators will chop up the data into segments that can be presented to identify a decision. Sometimes different human experts will see the same data segment and identify different decisions. Similarly, sometimes a human expert will examine one data segment and come up with two possible decisions, each of which are valid. INTEGRAT accounts for this situation in all stages of the creation and testing of the data and the AI robot.\n* Independent Claims 1 and 6 - Provide a friendly user interface to present data segments to the human expert and let them change the segmentation algorithm, or manually change a segment, if it could yield more than one decision.\n* Independent Claims 13, 17, and 20 – Provide a mechanism whereby multiple “correct” decisions can be assigned to one data segment, as well as AI training algorithms to support this ambiguity. Finally, allow the AI robot to come up with a single “best” decision when presented with ambiguous data, while also alerting users to “second best” decisions, and so on.\n","metadata":{}}]}