{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## This is the final notebook in a series that explained \"Reading the Robot Mind.\" I hope you'll find it insightful and interesting. \n\n### NOTE: This notebook uses the outputs of \"v16e-gpu-all-birdclef2023-mindreader\" (trained CNN) to just perform inference.\nBefore running, you must use the \"Add Data\" button, and add all the output of \"v16e gpu all birdclef2023 mindreader\" (another public notebook with the following direct link: https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader)\n\n#### Best Regards, \n#### ---Paul Alton Nussbaum","metadata":{}},{"cell_type":"markdown","source":"# Executive Summary\n* I coined the term \"Reading the Robot Mind\" to explain the process of presenting the internal workings of an Artificial Neural Network, or Artificial Intelligence (AI) system, in a way that is intuitive to Subject Matter Experts (SME). \n* The term \"explainability\" has been used extensively to describe methods that seek to explain how an AI system came to a particular output from a particular input, and can also be useful in troubleshooting.\n* Reading the robot mind goes further; allowing SME to see how much information is retained at each layer of the AI system presented in a format familiar to them - even if they are not programmers or are not familiar with the steps in the inference process.\n* Reading the robot mind also allows the SME to present hypothetical AI outputs and work backwards through the AI to create an approximation (exemplary) input.\n* In this case, SME are folks who can listen to birds or recordings of birds and determine which type of bird is making the sounds. The notebook therefore seeks to provide audio output as well as spectrographic images (both used by bird SME) to the user.\n* This is not new (see expired US Patent list at the end of the notebook) nor is it revolutionary. Indeed, this is intuitive for someone testing a microphone and data recording system. They do not want to look at long lists of numbers of recorded data. Instead, they will want to use that data to recreate an approximation of the audio, to listen to the recordings to see if they contain valid and useful information.\n* The notebook extends this intuitive and useful technique to individual neural network layers - working backwards towards a best estimate of the original input.\n* The user can even provide just the \"answer\" (select a bird at the final output layer of the AI solution), and the reading the robot mind system will work backwards through the entire automated process and AI layers to let the user hear a best approximation of what it has learned that bird sounds like.\n* You may be surprised at what you hear!","metadata":{}},{"cell_type":"markdown","source":"# This Notebook and Related Notebooks\n\nAs described above, a \"reading the robot mind\" system can be easily created by simultaneously training auto-encoders along each layer of the network (able to recreate the best approximation of the original input from the data available at that stage of the process). Many shy from this approach because of the added computational cost and time - as competition forces a race to market. I think it is a mistake to skip this step because the training cost is minimal compared to inference costs if the solution is successful. As an example, I cite the currently running Kagle competition where these autoencoders were not trained along with the generative solution, and now there is a scramble to \"read the robot mind,\" (https://www.kaggle.com/competitions/stable-diffusion-image-to-prompts) but there are many other examples, even with in-the-news stories about ChatGPT and the search for inputs that cause toxic or undesired responses.\n\nI have created these notebook to show that it is possible to make a system to read the robot mind, even when those autoencoders were not trained alongside the network. Below is a list of the notebooks related to this one:\n\n* https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader - This notebook focuses on the Segmentation and Feature Extraction aspects of the AI solution, allowing users to make modifications (Mel Scal Spectrum, number of resultant frequency bands, MFCC, number of resultant coefficients) and see and hear how much information is retained. It also allows the user to train a simple Convolutional Neural Network (CNN) AI solution and see and hear how much information is retained at each layer - using only a limited number of birds to speed up experimentation. Reading the robot mind learnings include selection of a feature exraction algorithm whereby SME can see and hear that important information has not been discarded, and can also see similarities in feature visualizations (spectrograms) for the same type of birds, while also spotting visual differences between different bird types. The final method used 5 second segments, each converted into an overlapping sequence of 32 Mel Scaled frequency bands.\n\n* https://www.kaggle.com/code/pnussbaum/v15h-all-birdclef2023-mindreader - This notebook allows the user to use their final decision related to segmentation and feature extraction, and convert and save the BirdClef2023 data into this format. It also allows a short amount of training of a CNN on this data; once again, allowing the SME to see and hear how much data is retained at each layer - this time using the entire data set. Learnings include identification of specific layers where important information seems to be discarded. For example, it was noted that the recreation of the input seemed to degrade significantly at certain layers of the network, and so these layers were specifically modified to improve the retention of important information (in the form of additional nerons for dense layers and additional filters for convolutional layers).\n\n* https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader - This notebook uses the final decisions noted above, and trains the entire CNN for a longer period of time, achieving better accuracy, and saving the trained AI system. There were several iterations of this notebook, with improvements made based on the reading the robot mind system, but also traditional troubleshooting techniques. An example of one of the traditional AI troubleshooting techniques was the identification of overfitting. This was remediated through the use of data augmentation (shifting the data in time by a small random value during training, etc.)\n\n* The current notebook you are reading - This notebook brings all of this together for the sake of the contest submission, as well as inference analysis and trouleshooting. The notebook lets you hear, visualize, and compare data every step and layer on the way. It allows comparison of samples from the same type of bird (so you can see and hear if they are similar) and different types of birds (so you can see and hear if they are different). Finally, this notebook provides the ability to work backwards through the system from a manually forced output, and let the user see and hear a best estimation of what the trained AI \"thinks\" that bird type sounds like. Learnings include the ability to hear elements of unique aspects particular to that bird type (short sequences of sound) all mashed together, with a good deal of extraneous noise mixed in.","metadata":{}},{"cell_type":"markdown","source":"## Loading the meta data","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ntrain_meta = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\ntotal_data_count = len(train_meta)\nprint(\"There are a total of \", total_data_count, \"  entries in the dataset\")\n# Uncomment the below two lines if you wish to see an example nmeta data entry.\n# print(\"Here is the first entry in the dataset metadata:\\n\")\n# print(train_meta.iloc[0])\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:52:09.843697Z","iopub.execute_input":"2023-05-02T20:52:09.844177Z","iopub.status.idle":"2023-05-02T20:52:09.971045Z","shell.execute_reply.started":"2023-05-02T20:52:09.844131Z","shell.execute_reply":"2023-05-02T20:52:09.969999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a list of all the possible classifications (labels)\nlabels = list(train_meta['primary_label'].unique())\nprint(\"There are \", len(labels), \" different labels in the whole dataset.\\n\")\nnum_birds = len(labels)\nepochs = 16\npre_trained = True\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:52:32.730358Z","iopub.execute_input":"2023-05-02T20:52:32.731274Z","iopub.status.idle":"2023-05-02T20:52:32.739269Z","shell.execute_reply.started":"2023-05-02T20:52:32.731229Z","shell.execute_reply":"2023-05-02T20:52:32.738053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_data = []\nfor label in labels : \n    image_data.append(train_meta[train_meta[\"primary_label\"] == label].iloc[:])\nimage_data = pd.concat(image_data).reset_index(drop=True) \ntotal_image_data_count = len(image_data)\nprint(\"There are \", total_image_data_count, \" different entries in the image data set.\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:52:37.173597Z","iopub.execute_input":"2023-05-02T20:52:37.174028Z","iopub.status.idle":"2023-05-02T20:52:37.732852Z","shell.execute_reply.started":"2023-05-02T20:52:37.173994Z","shell.execute_reply":"2023-05-02T20:52:37.731606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Helper functions to transform and untransform (reverse transform) and save resultant features","metadata":{}},{"cell_type":"code","source":"# These functions will help transform and untransform (reverse transformation) as well as\n# reducing the resolution of floating point feature values into 8-bit unsigned integers.\n# The goal is to convert each 5 second segment of auio into a 2D greyscale image of 8-bit uints.\n# After this is done, we can use these images as inputs to our classifier.\n\nimport librosa \nimport numpy as np\nfrom PIL import Image\nimport os\nimport soundfile as sf\n\n# These are the features to be extracted, as was decided from earlier notebook experiments\nuse_mfcc = False # otherwise use Mel Spectrogram\nn_mels = 32 # used for both Mel Spectrogram and MFCC\nn_coeff = 1 # number of cepstral coefficients (only used for MFCC)\n\n# Image dimensions\nif use_mfcc :\n    if n_coeff < n_mels :\n        num_rows = n_coeff \n    else :\n        num_rows = n_mels\nelse :\n    num_rows = n_mels\n    \nnum_columns = 313 # number of MFCC frames generated from 160000 samples of sudio (5 sec at 32000)\nnum_channels = 1\n\n# MAY BE A MISTAKE TO MAKE THIS A GLOBAL VALUE - but it's always forced to 32000\nsr = 32000 # default sampling rate if none provided\n\nWindow_Size = 5 # 5 seconds\n# Get number of samples for 5 seconds\nsegment = Window_Size * sr\n\n\ndef Audio_to_Domain(y, sr):\n    if use_mfcc:\n        feat = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=n_coeff, n_mels=n_mels, power=3) # raise to 3rd power before extracting coefficients to reduce noise\n    else :\n        feat = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=n_mels, power=1)\n    if feat.shape[1] <= num_columns:\n        pad_width = num_columns - feat.shape[1]\n        feat = np.pad(feat, pad_width=((0,0),(0,pad_width)), mode='constant')\n    return feat\n\ndef Domain_to_Audio(feat, sr):\n    if use_mfcc :\n        y = librosa.feature.inverse.mfcc_to_audio(feat, sr=sr, n_mels=n_mels, power=3)\n    else :\n        y = librosa.feature.inverse.mel_to_audio(feat, sr=sr, power=1) \n    return y\n\n\ndef Audio_Segment_Transform_and_Save(file_path, cat_name='dummy', is_save=True, train=True, sr=sr):\n\n    # First load the file\n    filename = file_path.replace(\"/\", \"_\")\n    file_path = \"/kaggle/input/birdclef-2023/train_audio/\" + file_path\n    audio, sr = librosa.load(file_path, sr = sr)\n\n    # Get number of samples for 5 seconds\n    buffer = segment\n\n    samples_total = len(audio)\n    samples_wrote = 0\n    counter = 1\n\n    # this will be a list of 5 second segments\n    feature_split = []\n    # each segment will have its own filename\n    feature_filenames = []\n    # each segment will also have its own min and range value for reconstruction (reverse transformation)\n    feature_mins = []\n    feature_ranges = []\n    while samples_wrote < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_wrote):\n            buffer = samples_total - samples_wrote\n\n        block = audio[samples_wrote : (samples_wrote + buffer)]\n        these_features = Audio_to_Domain(block, sr)\n        feature_split.append(these_features)\n\n        # Write Window_Size second segment\n        if is_save == True:\n            # had to add the letter \"z\" in front of my output files and paths, becuase contest submission searches output files alphabetically (!?)\n            out_filename = \"/kaggle/working/zimages/\" + cat_name + \\\n                \"/split_\" + str(counter) + \"_\" + filename + \".png\"\n            \n            feature_filenames.append(out_filename)\n            \n            # convert the features into an image for saving\n            \n            #scale to 0 - 255\n            fmax = these_features.max()\n            fmin = these_features.min()\n            frange = fmax - fmin\n            # these_features = 255 * ((these_features - fmin)/frange)\n            these_features = np.array((((these_features - fmin) / frange)*255), dtype='uint8')\n            # save the min and range for later reconstruction (reverse transform, or un-transform) back to audio\n            feature_mins.append(fmin)\n            feature_ranges.append(frange)\n            this_image = Image.fromarray(these_features, mode = 'L')\n            # make the directory if it doesn't yet exist\n            if not os.path.exists(os.path.dirname(out_filename)):\n                try:\n                    os.makedirs(os.path.dirname(out_filename))\n                except OSError as exc: # Guard against race condition\n                    if exc.errno != errno.EEXIST:\n                        raise            \n            # save the file\n            this_image.save(out_filename, format=\"PNG\")\n            \n        counter += 1\n        samples_wrote += buffer\n    return feature_split, sr, feature_filenames, feature_mins, feature_ranges\n\ndef Load_and_UnTransform(filename):\n    this_image = Image.open(filename)\n    block = Domain_to_Audio(this_image, sr)\n    return block, sr\n    \n\ndef Save_Features(_df, train=True):\n    data = []\n    \n    # When importing directories, we use the \"input\" folder instead of the \"working\" folder\n\n    if os.path.exists(os.path.dirname(\"/kaggle/input/v16e-gpu-all-birdclef2023-mindreader/zimages/\")):  # was \"working\" instead of \"input/v16e-gpu-all-birdclef2023-mindreader\"\n        print(\"WARNING: Image files already saved. Aborting save to avoid corrupted image databse.\")\n        data_df = pd.DataFrame(data, columns=['primary_label', 'original_filename', 'filename', 'fmin', 'frange'])\n        # had to add the letter z in front of filenames and paths so competition submission can find the output file (alphabetically)\n        data_df = pd.read_csv(\"/kaggle/input/v16e-gpu-all-birdclef2023-mindreader/zimages.csv\")  # was \"working\" instead of \"input/v16e-gpu-all-birdclef2023-mindreader\"\n    else:\n        for index, row in _df.iterrows(): # was tqdm(_df.iterrows()):\n            audio_lst, sr, filenames, mins, ranges = Audio_Segment_Transform_and_Save(row[\"filename\"], cat_name = row[\"primary_label\"], is_save = True, train = train)\n            for idx, y in enumerate(audio_lst):\n                data.append([row[\"primary_label\"], row[\"filename\"], filenames[idx], mins[idx], ranges[idx]])\n    \n        data_df = pd.DataFrame(data, columns=['primary_label', 'original_filename', 'filename', 'fmin', 'frange'])\n        data_df.to_csv(\"/kaggle/working/zimages.csv\", index=False)\n\n    return data_df\n    \n    \n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:52:48.833731Z","iopub.execute_input":"2023-05-02T20:52:48.834164Z","iopub.status.idle":"2023-05-02T20:52:48.902034Z","shell.execute_reply.started":"2023-05-02T20:52:48.834119Z","shell.execute_reply":"2023-05-02T20:52:48.900243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# See and Hear the Feature Extraction\n#### Try a number of Feature Extraction Algorithms and test the following:\n* Do the features visually look similar for the same bird, and different for different birds?\n* Is the similarity/difference enough to be able to visually classify which bird is which?\n* What about the audio - is the recreated sound clear enough for an expert to identify the bird?\n* If the answer is \"no\" to any of the above, allow fine tuning by the user (three different feature extractions are compared below)","metadata":{}},{"cell_type":"code","source":"# Visually examine feature extraction algorithms\nimport IPython\nimport matplotlib.pyplot as plt\n\ndef Compare_Feature_Extraction(image_data, list_to_compare) :\n    global sr, segment\n    list_len = len(list_to_compare)\n    if list_len < 2 :\n        print(\"Must provide at least 2 items - cannot compare.\\n\")\n        return()\n    else:\n        plt.figure(figsize=(10, 10))\n        for i in range(list_len):\n            # grab from the image set\n            audio_filename = image_data.at[list_to_compare[i],\"filename\"]\n            common_name = image_data.at[list_to_compare[i],\"common_name\"]\n            # load the audio data\n            audio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n            # Take first 5 second \"segment\"\n            audio = audio[0:segment]\n            # Extract the feaures\n            feat = Audio_to_Domain(audio, sr)\n            # Move the range from current min and max, into 0 to 255 8 bit integers\n            fmin = feat.min()\n            fmax = feat.max()\n            frange = fmax - fmin\n            feat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\n            # Plot the features\n            ax = plt.subplot(3, 3, i + 1)\n            plt.imshow(feat, cmap='gray', aspect=(num_columns/num_rows), interpolation = 'None')\n#            plt.imshow(images[i].numpy().astype(\"uint8\"), cmap='gray', aspect=(num_columns/num_rows), interpolation = 'None')\n            plt.title(common_name)\n            plt.axis(\"off\")\n    return()\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:52:59.963356Z","iopub.execute_input":"2023-05-02T20:52:59.963836Z","iopub.status.idle":"2023-05-02T20:52:59.975604Z","shell.execute_reply.started":"2023-05-02T20:52:59.96379Z","shell.execute_reply":"2023-05-02T20:52:59.974137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Compare features from the SAME birds (first 5 sec of audio used)\\n\")            \n# Compare the first 10 entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 1, 2, 3, 4, 5, 6, 7, 8])","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:53:20.704557Z","iopub.execute_input":"2023-05-02T20:53:20.705248Z","iopub.status.idle":"2023-05-02T20:53:34.962259Z","shell.execute_reply.started":"2023-05-02T20:53:20.705204Z","shell.execute_reply":"2023-05-02T20:53:34.960899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from the same bird\n* Comparing features from the same bird now yields more similarities\n* Seems like now we have plenty of information (maybe too much? Will it slow down our training algorithm or cause it to overfit?)","metadata":{}},{"cell_type":"code","source":"print(\"Compare features from the DIFFERENT birds (first 5 sec of audio used)\\n\")            \n# Compare the first 10 entries in the image_data\nCompare_Feature_Extraction(image_data, [0, 50, 150, 200, 225, 350, 430])","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:53:34.964705Z","iopub.execute_input":"2023-05-02T20:53:34.965281Z","iopub.status.idle":"2023-05-02T20:53:35.986899Z","shell.execute_reply.started":"2023-05-02T20:53:34.965245Z","shell.execute_reply":"2023-05-02T20:53:35.985488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The above shows extracted features for sounds samples from DIFFERENT birds\n* Different birds indeed seem to have different features, but also instantaneous similarities (which may be differentiated as they change over time). \n* \"Different versus similar\" birds start to make itself more visible, but with regard to that time window, it looks like 1.5 to 2 seconds of changing notes might be needed to differentiate birds (since there are 313 samples in a 5 second window, this equates to 95 to 125 feature vectors.\n* Still not sure if a bird expert would use these images to classify birds.","metadata":{}},{"cell_type":"markdown","source":"## Let's listen to the untransformation (recreate audio from the extracted features)\n* We'll only listen to the first segment from the first bird","metadata":{}},{"cell_type":"code","source":"audio_filename = image_data.at[0,\"filename\"]\ncommon_name = image_data.at[0,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n# Extract the feaures\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\n# Go back to the original range of values\nfeat = ((feat/255)*frange)+fmin\n# Recreate the audio from the extracted features\nrecreated_audio = Domain_to_Audio(feat, sr)\n\nprint(\"In BLUE is the original audio of the first segment of the first audio file\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"Superimposed in ORANGE the RECREATED (un-transformed) audio of the first segment of the first image file\\n\")\nlibrosa.display.waveshow(recreated_audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:53:35.988666Z","iopub.execute_input":"2023-05-02T20:53:35.990033Z","iopub.status.idle":"2023-05-02T20:53:39.775503Z","shell.execute_reply.started":"2023-05-02T20:53:35.989971Z","shell.execute_reply":"2023-05-02T20:53:39.774261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"This is the audio playback of the RECREATED audio file\")\nIPython.display.Audio(data = recreated_audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:53:39.777683Z","iopub.execute_input":"2023-05-02T20:53:39.778034Z","iopub.status.idle":"2023-05-02T20:53:39.795879Z","shell.execute_reply.started":"2023-05-02T20:53:39.778Z","shell.execute_reply":"2023-05-02T20:53:39.794469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above lets us listen to recreation of the original input from the Extracted Features## The above showed\n* The audio is definitely more clear\n* It seems obvious that \"a human expert will be able to use the RECREATED audio to correctly classify birds.\"\n* Naturally, if they couuld perform the classification with the raw audio\n* Segmentation is still the open question - if the 5 second window does not containt enough of that bird's call, etc...","metadata":{}},{"cell_type":"markdown","source":"# Build Convolutional Neural Network\n## Build a simple CNN to perform classification\n* Use sequences of Conv2D and MaxPool to find edges, textures, and patterns\n* NOTE: You can change Conv2D filter counts and shapes, as well as MaxPool2D shapes, but if you add or remove layers, you'll need to modify the backwards process (manually coded later in the notebook) as I have not yet implemented a method that figures out the layer structure automatically.","metadata":{}},{"cell_type":"code","source":"from tensorflow import keras\nfrom keras.models import Model\nfrom keras.layers import Dense,Conv2D,Flatten,MaxPool2D,Concatenate, Activation\nfrom keras.layers import Dropout,BatchNormalization, Input, Rescaling, RandomTranslation, GaussianDropout\n\ninputs = Input(shape = (num_rows, num_columns, 1))\n\n# Rescale the input (arrives as 0-255 8-bit uint grayscale pixels)\nmodel = Rescaling(1./255)(inputs)\n### Added image augmentation by shifting input left or right by 20% (1 second) and filling with 0's\nmodel = RandomTranslation(0, .2, fill_mode=\"constant\", fill_value=0.0,)(model)\n### Added multiplicative (10%) 1-centered gaussian noise for data augmentation\nmodel = GaussianDropout(0.1)(model)\n\n# This model has a sequential set of Conv2D and MaxPool layers finished with a 32 Dense \n# In this way, the model seeks to embed the entire image into a 32 dim vector\nmodel1 = Conv2D(filters=64, kernel_size=(3, 3), padding='SAME', activation='relu')(model)   # was 32 filters\nmodel1 = Conv2D(filters=128, kernel_size=(3, 3), padding='SAME', activation='relu')(model1)  # was 32 filters\nmodel1 = MaxPool2D(pool_size=(2, 2))(model1)\n\nmodel1a = Conv2D(filters=128, kernel_size=(5, 5), padding='SAME', activation='relu')(model1) # was 64 filters\nmodel1a = Dropout(0.2)(model1a)\nmodel1a = Conv2D(filters=64, kernel_size=(3, 3), padding='SAME', activation='relu')(model1a)# was 32 filters\nmodel1a = Dropout(0.2)(model1a)\n\nmodel1b = Conv2D(filters=64, kernel_size=(3, 3), padding='SAME', activation='relu')(model1a)# was 32 filters\nmodel1b = Conv2D(filters=128, kernel_size=(3, 3), padding='SAME', activation='relu')(model1b)# was 32 filters\n# Lets compress the time down\nmodel1b = MaxPool2D(pool_size=(2, 2))(model1b)\n\nmodel1c = Conv2D(filters=256, kernel_size=(5, 5), padding='SAME', activation='relu')(model1b)# was 64 filters\n\n# Flatten and use Dense to find a lower dim embedding\nmodel1c = Flatten()(model1c)\nmodel1c = Dense(128, activation = \"relu\")(model1c)\n\nbird_classification = Dense(num_birds, activation = 'softmax')(model1c)\n\nmodel = Model(inputs=inputs, outputs=bird_classification)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:55:41.695635Z","iopub.execute_input":"2023-05-02T20:55:41.696034Z","iopub.status.idle":"2023-05-02T20:55:43.12823Z","shell.execute_reply.started":"2023-05-02T20:55:41.695998Z","shell.execute_reply":"2023-05-02T20:55:43.127234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"###########################################################################################\n# If the user changes the layers of the network, below parameters must also be changed    #\n###########################################################################################\n# NOTE This is manually configured. Need to make this automated.\nnum_conv = 7 # number of convolutional layers\nconv_layer = [4, 5, 7, 9, 11, 12, 14] # the layer (counting from layer 0, the input layer) of the Conv2D's'\nflatten_layer = 15\nembed_layer = 16   # The first Dense layer after the Flatten\nclass_layer = 17   # The last Dense layer (with Softmax Activation) that determines the final classification\nscale_size = [1, 1, 2, 1, 1, 2, 1] # any scaling (due to strides or max pooling) that took place just prior to or at that conv layer (not cumulative)\n###########################################################################################\n# If the user changes the layers of the network, above parameters must also be changed    #\n###########################################################################################\n\n#######################################################################################################################\n# If changing the shape of the last convolutional later or embed layer, please also change this next lines            #\n# Sorry this is hard coded. Need to make it autoconfigured.                                                           #\n#######################################################################################################################\npre_flatten = (8, 78, 256, 128)   \ninto_flatten = (8, 78, 256)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:55:47.888621Z","iopub.execute_input":"2023-05-02T20:55:47.889064Z","iopub.status.idle":"2023-05-02T20:55:47.89678Z","shell.execute_reply.started":"2023-05-02T20:55:47.889024Z","shell.execute_reply":"2023-05-02T20:55:47.895272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Here you can see the Model Summary","metadata":{}},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:55:52.299123Z","iopub.execute_input":"2023-05-02T20:55:52.299561Z","iopub.status.idle":"2023-05-02T20:55:52.354188Z","shell.execute_reply.started":"2023-05-02T20:55:52.299522Z","shell.execute_reply":"2023-05-02T20:55:52.352889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Here you can see the architecture of the model","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\ntf.keras.utils.plot_model(model)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:55:56.645268Z","iopub.execute_input":"2023-05-02T20:55:56.645703Z","iopub.status.idle":"2023-05-02T20:55:56.857218Z","shell.execute_reply.started":"2023-05-02T20:55:56.645662Z","shell.execute_reply":"2023-05-02T20:55:56.856077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the Trained CNN\n","metadata":{}},{"cell_type":"code","source":"if os.path.exists(os.path.dirname(\"/kaggle/input/v16e-gpu-all-birdclef2023-mindreader/zclassifier/\")):  # was \"working\" instead of \"input/v16e-gpu-all-birdclef2023-mindreader\"\n    print(\"WARNING: Trained CNN already saved. Aborting training to avoid corrupted CNN.\")\n    checkpoint_model_path = \"/kaggle/input/v16e-gpu-all-birdclef2023-mindreader/zclassifier\"\n    model = keras.models.load_model(checkpoint_model_path)\n    pre_trained = True\nelse:\n    print(\" ERROR - Missing trained CNN weights. You must ADD DATA from v16e-gpu-all-birclef2023-mindreader before running this notebook.\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:57:12.009764Z","iopub.execute_input":"2023-05-02T20:57:12.011129Z","iopub.status.idle":"2023-05-02T20:57:16.124215Z","shell.execute_reply.started":"2023-05-02T20:57:12.011081Z","shell.execute_reply":"2023-05-02T20:57:16.122979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Examine the CNN Filters\n## Visualize each layer of the trained CNN\n* We wish to examine the trained CNN model. In particular, we will be looking to see if we have too many (redundant) convolutional filters, or if there are other issues that can be discerened by looking into the trained CNN.\n* We will do this by traversing the internal workings of the trained CNN weights.","metadata":{}},{"cell_type":"code","source":"model_explore = [] # list of models that stop at each convolutional layer\nConv2D_weights = [] # list of filters for convolutional layer\nConv2D_biases = [] # list of filter biases for convolutional layer\n\nfor i in range (num_conv):\n    conv_ind = conv_layer[i]\n    print(\"Conv2D \", i, \"occurs at layer\",conv_ind)\n    model_explore.append(Model(inputs=model.inputs, outputs=model.get_layer(index=conv_ind).output))\n    print(\"Including prior layers, will get an Input Shape of:\",model_explore[i].input_shape)\n    print(\"and have an Output Shape of:\",model_explore[i].output_shape )\n    Conv2D_weights.append(np.copy(model.get_layer(index=conv_ind).get_weights()[0]))\n    Conv2D_biases.append(np.copy(model.get_layer(index=conv_ind).get_weights()[1]))\n    # These are the filters\n    print(\"Final Conv2D has the following rows, columns, color depth, number of filters\", Conv2D_weights[i].shape)\n    # These are the biases\n    print(\"and bias for each filter\", Conv2D_biases[i].shape)\n    print(\"------------------------\")\n\nmodel_flatten = Model(inputs=model.inputs, outputs=model.get_layer(index=flatten_layer).output)\nmodel_embed = Model(inputs=model.inputs, outputs=model.get_layer(index=embed_layer).output)\nmodel_class = Model(inputs=model.inputs, outputs=model.get_layer(index=class_layer).output)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:58:02.730995Z","iopub.execute_input":"2023-05-02T20:58:02.732049Z","iopub.status.idle":"2023-05-02T20:58:02.801432Z","shell.execute_reply.started":"2023-05-02T20:58:02.732003Z","shell.execute_reply":"2023-05-02T20:58:02.800135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create multiple sub-CNN networks\n* Each network will be identical to the original trained network, except the output layer will be one of the Conv2D layers\n* We want to input an image (in this case, features extracted from audio bird calls, transformed into an 8 bit resolution grayscale image) and output the results of that last Conv2D layer.\n* From this intermediate output, we will work backwards through the CNN and attempt to recreate the original image (later in this notebook).\n* We will also use these sub-networks to visualize the CNN filters as described above.","metadata":{}},{"cell_type":"code","source":"# PANv12f now we use a list of models that can be traversed.\nfor conv_l in range(num_conv):\n    print(\"Overall Explore \",conv_l,\" Model Summary:\", model_explore[conv_l].summary(), \"\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:58:11.82934Z","iopub.execute_input":"2023-05-02T20:58:11.829759Z","iopub.status.idle":"2023-05-02T20:58:12.089084Z","shell.execute_reply.started":"2023-05-02T20:58:11.829725Z","shell.execute_reply":"2023-05-02T20:58:12.087856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's observe the internal models of the CNN visually\n* Each convolution filter has weights that form the pattern \"it is looking for.\" \n* More specifically, image areas that most closely match the weights of the convolutional filter will yield a higher numerical value when convolved (multiplied) by that filter.\n* When MaxPooling, it is these most closely matching parts of the image that will have maximal values, and therefore remain, while other, less closely matching portions of the image will be thrown away since they will vield lower convolution values.\n* In this way we can visualize the types of image pixel groupings each CNN filter \"is looking for.\"\n* When multiple CNN layers are connected in sequence, the last layer receives a CNN value from multiple filters in the prior layer. Here too, the last CNN layer is \"looking for\" combinations of prior layer CNN filter convolution results. We can literally combine those prior CNN layer visualizations to form a visualization for each of the last CNN layer filters.\n* In this way, we build up visualizations of what each filter of each CNN layer \"is looking for.\" \n* Layers like Droput are ignored in this analysis\n* Layers like MaxPooling are handled with a scaling factor\n### NOTE: Similar to the \"expansion\" method shown here https://distill.pub/2020/circuits/visualizing-weights/","metadata":{}},{"cell_type":"code","source":"# a list of filter patches, one for each convolutional layer\nnum_filt = []            # number of filters in each Conv2D layer\nfilt_siz = []            # the size of each filter (assume they are square for now)\nborder_siz = []          # the size of the border around the center pixel of the square (assume filters are odd numbered dimensions like 3, 5, 7, etc.)\ncumulative_border = []   # The cumulative effect of the borders from prior Conv3D layers\nfilter_patches = []      # The \"mind reading\" visualization of each filter - equal to the filter at the first Conv2D layer, but more complex later\n\nfor i in range (num_conv):\n    conv_ind = conv_layer[i]\n    print(\"Show the reconstructed image patch for each filter of convolutional layer \", i, \"which occured at layer \",conv_ind, \"in the original model\")\n    num_filt.append(Conv2D_weights[i].shape[3])\n    # filter Size\n    filt_siz.append(Conv2D_weights[i].shape[0])\n    # Border size\n    border_siz.append(int(filt_siz[i] / 2))\n    cb = 0\n    if i == 0 : \n        cumulative_border.append(0)\n    else:\n        for j in range(i) :\n            cb += border_siz[j]\n        cumulative_border.append(cb)\n    print(\"Conv Layer\", i, \" filter Size:\", filt_siz[i], \" x \", filt_siz[i], \n          \" and border size (padding) of \", border_siz[i], \" and cumulative border from prior layers of \", cumulative_border[i])\n    filter_patches.append(np.zeros((filt_siz[i] + 2 * cumulative_border[i], \n                                    filt_siz[i] + 2 * cumulative_border[i], \n                                    num_filt[i])))\n\n    # For now, draw eight filters on a row for display purposes\n    this_num_filt = num_filt[i]\n    if (this_num_filt % 8 == 0) :\n        drows = int(this_num_filt/8)\n    else :\n        drows = int(this_num_filt/8) + 1\n    dcols = 8\n    width = 24\n    height = int(this_num_filt / 4)\n    fig, ax = plt.subplots(nrows=drows, ncols=dcols, figsize=(width, height))\n    \n    thisrow = 0\n    thiscol = 0\n    # temporary variables for filter size and cumulative border size\n    cb = cumulative_border[i]\n    siz = filt_siz[i]\n    image_patch = np.zeros((siz + 2 * cb, siz + 2 * cb))\n    for this_filter in range(this_num_filt) :\n        if i == 0 :\n            image_patch = np.copy(Conv2D_weights[i][:, :, 0, this_filter])\n        else:\n            image_patch.fill(0)\n            for x in range(cb, siz + cb, 1) :\n                for y in range(cb, siz + cb, 1) :\n                    for depth in range(num_filt[i-1]) :\n                        image_patch[x-cb:x+cb+1, y-cb:y+cb+1] += np.copy((Conv2D_weights[i][x-cb, y-cb, depth, this_filter] * \n                                                                         filter_patches[i-1][:,:,depth]))\n    \n        filter_patches[i][:, :, this_filter] = np.copy(np.clip(image_patch, 0, 1e999)) # remove negative numbers (due to \"relu\" activation we used)\n        \n        # display the filter \n        if (drows > 1) :\n            ax[thisrow, thiscol].imshow(filter_patches[i][:, :, this_filter], cmap=\"gray\")\n        else :\n            ax[thiscol].imshow(filter_patches[i][:, :, this_filter], cmap=\"gray\")\n        thiscol += 1\n        if (thiscol >=8) :\n            thisrow += 1\n            thiscol = 0\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T20:58:20.346702Z","iopub.execute_input":"2023-05-02T20:58:20.347651Z","iopub.status.idle":"2023-05-02T20:59:41.584547Z","shell.execute_reply.started":"2023-05-02T20:58:20.3476Z","shell.execute_reply":"2023-05-02T20:59:41.583614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above shows the filters of each Conv2D layer, multiplied by prior filters, to create a cumulative pattern being targeted by each successive layer (top to bottom)","metadata":{}},{"cell_type":"markdown","source":"## What can we learn from the above?\n* We can see that successive convolution layers manage succesive levels of complexity. These are often described as edge detection first, then these are combined to form textures. Finally, the last convolution layer combines these textures to form patterns it is looking for in order to correctly identify birds. \n* Do we have too many convolutional layers in the middle? Examining the filters, are all the middle layers \"adding value?\"\n* Looking at the last layer (the last group of images) we can see patterns the CNN is looking for. Specifically, look for the white areas and dark areas, and see if the white area is sloped upwards (the bird is going from a high frequency to a low frequency) or downwards (bird goes from low frequency to high frequency, otherwise known as a \"chirp\").\n* Notice also that the specific frequency (musical notes) that the bird is singing is NOT shown at the convolution layer, just the pattern, which might be found at low frequencies or high ones. We would need to reconstruct the signal from the convolutional layer (work backwards through the network) to be able to discern what frequency / time does the pattern occur.\n","metadata":{}},{"cell_type":"markdown","source":"# What Does Bird \"X\" Sound Like?\n## What if we don't have an input signal, and just work backwards from a classification (forced output)?\n* We will repeat this for the first few birds, to see what the \"Robot Mind\" thinks that particular bird sounds like\n","metadata":{}},{"cell_type":"markdown","source":"## Starting with bird class 0 - \"abethr1\"\n* We'll hear what an example original audio sounds like, and see the input features associated with it\n* After that, we'll work backwards through the NN to approximate an input feature set (image) and from that audio best approximation\n* Finally, we'll listen to what the NN has learned each bird \"sounds like\" - you may be surprised!","metadata":{}},{"cell_type":"markdown","source":"### Listen to an example of what that bird sounds like","metadata":{}},{"cell_type":"code","source":"bird_class = 0\nexample_audio_index = 0\naudio_filename = image_data.at[0,\"filename\"]\n# Image files are now in the \"input/v16e-gpu-all-birdclef2023-mindreader\" directory instead of \"working\"\naudio_filename = audio_filename.replace(\"working\", \"input/v16e-gpu-all-birdclef2023-mindreader\")\ncommon_name = image_data.at[0,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n\nprint(\"Here is an example of original audio of this \", common_name, \" bird sounds like\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:03:49.768313Z","iopub.execute_input":"2023-05-02T22:03:49.768818Z","iopub.status.idle":"2023-05-02T22:03:50.301431Z","shell.execute_reply.started":"2023-05-02T22:03:49.768779Z","shell.execute_reply":"2023-05-02T22:03:50.299713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### For the above example, show the extracted features","metadata":{}},{"cell_type":"code","source":"# These are the features for the above audio sample\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\nprint(\"These are teh features that will be used as input for the above example audio.\\n\")\nplt.imshow(feat, cmap='gray', aspect=int(num_columns/num_rows) , interpolation = 'None')\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:03:50.303979Z","iopub.execute_input":"2023-05-02T22:03:50.305343Z","iopub.status.idle":"2023-05-02T22:03:50.539873Z","shell.execute_reply.started":"2023-05-02T22:03:50.305283Z","shell.execute_reply":"2023-05-02T22:03:50.538048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Force the output category of the NN to be the above bird","metadata":{}},{"cell_type":"code","source":"# Let's see if we can work backwards from a classification - in this case, we'll use the frst bird\nbird_class = 0 # You can pick a different bird here if you wish\nforce_class = np.zeros(num_birds)\nforce_class[bird_class] = 1\n\nprint(\"Here is the forced output of our classification layer (the last Dense layer with Softmax activation)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_birds)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, force_class)\n\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:03:50.542275Z","iopub.execute_input":"2023-05-02T22:03:50.543302Z","iopub.status.idle":"2023-05-02T22:03:51.155212Z","shell.execute_reply.started":"2023-05-02T22:03:50.54324Z","shell.execute_reply":"2023-05-02T22:03:51.153746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's calculate the Embed layer for this forced classification\n# Weights from the embed layer to the Classification layer\nclass_weights = np.copy(model_class.get_layer(index=class_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", class_weights.shape)\n\nnum_embed_neurons = class_weights.shape[0]\nnum_class_neurons = class_weights.shape[1]\nembed_data = np.zeros(num_embed_neurons)\nfor i in range(num_embed_neurons):\n    for j in range(num_class_neurons):\n        embed_data[i] += force_class[j] * class_weights[i, j]\n\nprint(\"Here is the forced output of our embed layer (the first Dense layer after flattened last Conv layer)\")\n\n# Creating bar chart of Embed layer output values\nx = np.arange(num_embed_neurons)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, embed_data)\n\n# Show plot\nplt.show()\n\n# Weights from the embed layer to the Flatten-ed last Conv layer\nembed_weights = np.copy(model_embed.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the flattened last Conv layer to the embed has the shape \", embed_weights.shape)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:03:51.158651Z","iopub.execute_input":"2023-05-02T22:03:51.159245Z","iopub.status.idle":"2023-05-02T22:03:51.655211Z","shell.execute_reply.started":"2023-05-02T22:03:51.159193Z","shell.execute_reply":"2023-05-02T22:03:51.653645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Work backwards through the NN to recreate a best approximation of the input features, and audio, of the above bird.\n* Essentially \"What has the NN learned to be the audio feature set for the above bird, and what does it sound like?\"","metadata":{}},{"cell_type":"code","source":"# Let's calculate the last Conv2D layer for this forced classification\n# Weights from the last Conv2D layer to the embed layer\nlast_conv_weights_flat = np.copy(model.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", last_conv_weights_flat.shape)\n\nlast_conv_weights = np.copy(np.reshape(last_conv_weights_flat, pre_flatten))\nprint(\"Weights array shape from the embed layer after undoing the Flatten \", last_conv_weights.shape)\n\n# The last convolutional network  is a 3 dimensional arrax of dimensions rows x columns x filters\nnum_conv_rows = last_conv_weights.shape[0]\nnum_conv_cols = last_conv_weights.shape[1]\nnum_conv_filt = last_conv_weights.shape[2]\n# In addition, each of these rows x columns x filter outputs is connected to a dense layer with a number of neurons (called the embed layer)\nnum_embed_neurons = last_conv_weights.shape[3]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer (called embed_data here) multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor i in range(num_embed_neurons):\n    embed_value = embed_data[i] # using local variables to speed things up\n    for j in range(num_conv_filt):\n        this_filter_patch = last_filter_patches[:,:,j]\n        this_conv_to_embed_weights = last_conv_weights[:,:,j,i]\n        for k in range(border, num_conv_cols + border, 1):\n            for l in range(border, num_conv_rows + border, 1):\n                conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * this_conv_to_embed_weights[l-border,k-border] * embed_value\n                \n\nconv_data2 = np.copy(conv_data1)\nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"Here is the best approximation of the input for this bird\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:03:51.657762Z","iopub.execute_input":"2023-05-02T22:03:51.659084Z","iopub.status.idle":"2023-05-02T22:04:19.671927Z","shell.execute_reply.started":"2023-05-02T22:03:51.659022Z","shell.execute_reply":"2023-05-02T22:04:19.668436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = model.predict(np.reshape(conv_data, (1, num_rows, num_columns, 1)))\n\nprint(\"As noted above, recreation of input does not mean that recreation will be classified the same way. \\n\",\n      \"In fact, using the above as input, the NN categorizes this as bird number \", x.argmax())","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.673638Z","iopub.status.idle":"2023-05-02T22:04:19.674191Z","shell.execute_reply.started":"2023-05-02T22:04:19.673938Z","shell.execute_reply":"2023-05-02T22:04:19.673967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", conv_data.min(), conv_data.max())\nuntrans_1 = Domain_to_Audio(conv_data, sr)\nprint(\"This is the recreated audio, (un-transformed) hear and see as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)\nIPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.675952Z","iopub.status.idle":"2023-05-02T22:04:19.676495Z","shell.execute_reply.started":"2023-05-02T22:04:19.676216Z","shell.execute_reply":"2023-05-02T22:04:19.676244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above is what the AI \"thinks bird class 0 - abethr1 - sounds like.\" \nIf you'll forgive the personification, pretending like the AI is actually thinking\n### What can we learn from the above?\n* We can see that certain important frequencies appear in the same position (in the extracted feature image)\n* We can see that certain important cadences (low to high notes, versus high to low notes) can be spotted\n* Some information, such as rhythm/timing of repeated sounds, seems to be completely lost.\n* We see that quite a bit of information has been thrown away (I can't tell what bird it is from the audio, and probabbly neither can an expert).\n* Could this discarded information be important? Experts would say \"yes\" but AI researchers migth be OK with a good score only.","metadata":{}},{"cell_type":"markdown","source":"## Next is bird class 1 - \"abhori1\"\n* We'll hear what an example original audio sounds like\n* After that, we'll work backwards through the NN to approximate an input feature set (image) and from that audio best approximation\n* Listen to what the NN has learned each bird \"sounds like\" - you may be surprised!","metadata":{}},{"cell_type":"markdown","source":"### Listen to an example of what that bird sounds like","metadata":{}},{"cell_type":"code","source":"bird_class = 1\nexample_audio_index = 15\naudio_filename = image_data.at[example_audio_index,\"filename\"]\n# Image files are now in the \"input/v16e-gpu-all-birdclef2023-mindreader\" directory instead of \"working\"\naudio_filename = audio_filename.replace(\"working\", \"input/v16e-gpu-all-birdclef2023-mindreader\")\ncommon_name = image_data.at[example_audio_index,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n\nprint(\"Here is an example of original audio of this \", common_name, \" bird sounds like\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.678774Z","iopub.status.idle":"2023-05-02T22:04:19.679332Z","shell.execute_reply.started":"2023-05-02T22:04:19.679086Z","shell.execute_reply":"2023-05-02T22:04:19.679114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### For the above example, show the extracted features","metadata":{}},{"cell_type":"code","source":"# These are the features for the above audio sample\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\nprint(\"These are the features that will be used as input for the above example audio.\\n\")\nplt.imshow(feat, cmap='gray', aspect=int(num_columns/num_rows) , interpolation = 'None')\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.681252Z","iopub.status.idle":"2023-05-02T22:04:19.681801Z","shell.execute_reply.started":"2023-05-02T22:04:19.681569Z","shell.execute_reply":"2023-05-02T22:04:19.681595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Force the output category of the NN to be the above bird","metadata":{}},{"cell_type":"code","source":"# Let's see if we can work backwards from a classification - in this case, we'll use the frst bird\nbird_class = 1 # You can pick a different bird here if you wish\nforce_class = np.zeros(num_birds)\nforce_class[bird_class] = 1\n\nprint(\"Here is the forced output of our classification layer (the last Dense layer with Softmax activation)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_birds)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, force_class)\n\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.683126Z","iopub.status.idle":"2023-05-02T22:04:19.683627Z","shell.execute_reply.started":"2023-05-02T22:04:19.683372Z","shell.execute_reply":"2023-05-02T22:04:19.683418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's calculate the Embed layer for this forced classification\n# Weights from the embed layer to the Classification layer\nclass_weights = np.copy(model_class.get_layer(index=class_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", class_weights.shape)\n\nnum_embed_neurons = class_weights.shape[0]\nnum_class_neurons = class_weights.shape[1]\nembed_data = np.zeros(num_embed_neurons)\nfor i in range(num_embed_neurons):\n    for j in range(num_class_neurons):\n        embed_data[i] += force_class[j] * class_weights[i, j]\n\nprint(\"Here is the forced output of our embed layer (the first Dense layer after flattened last Conv layer)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_embed_neurons)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, embed_data)\n\n# Show plot\nplt.show()\n\n# Weights from the embed layer to the Flatten-ed last Conv layer\nembed_weights = np.copy(model_embed.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the flattened last Conv layer to the embed has the shape \", embed_weights.shape)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.685343Z","iopub.status.idle":"2023-05-02T22:04:19.685991Z","shell.execute_reply.started":"2023-05-02T22:04:19.68575Z","shell.execute_reply":"2023-05-02T22:04:19.685777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Work backwards through the NN to recreate a best approximation of the input features, and audio, of the above bird.\n* Essentially \"What has the NN learned to be the audio feature set for the above bird, and what does it sound like?\"","metadata":{}},{"cell_type":"code","source":"# Let's calculate the last Conv2D layer for this forced classification\n# Weights from the last Conv2D layer to the embed layer\nlast_conv_weights_flat = np.copy(model.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", last_conv_weights_flat.shape)\n\nlast_conv_weights = np.copy(np.reshape(last_conv_weights_flat, pre_flatten))\nprint(\"Weights array shape from the embed layer after undoing the Flatten \", last_conv_weights.shape)\n\n# The last convolutional network  is a 3 dimensional arrax of dimensions rows x columns x filters\nnum_conv_rows = last_conv_weights.shape[0]\nnum_conv_cols = last_conv_weights.shape[1]\nnum_conv_filt = last_conv_weights.shape[2]\n# In addition, each of these rows x columns x filter outputs is connected to a dense layer with a number of neurons (called the embed layer)\nnum_embed_neurons = last_conv_weights.shape[3]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer (called embed_data here) multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor i in range(num_embed_neurons):\n    embed_value = embed_data[i] # using local variables to speed things up\n    for j in range(num_conv_filt):\n        this_filter_patch = last_filter_patches[:,:,j]\n        this_conv_to_embed_weights = last_conv_weights[:,:,j,i]\n        for k in range(border, num_conv_cols + border, 1):\n            for l in range(border, num_conv_rows + border, 1):\n                conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * this_conv_to_embed_weights[l-border,k-border] * embed_value\n                \nconv_data2 = np.copy(conv_data1) # [border:num_conv_rows + border, border:num_conv_cols + border])\nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"Here is the best approximation of the input for this bird\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.688232Z","iopub.status.idle":"2023-05-02T22:04:19.689335Z","shell.execute_reply.started":"2023-05-02T22:04:19.689061Z","shell.execute_reply":"2023-05-02T22:04:19.6891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = model.predict(np.reshape(conv_data, (1, num_rows, num_columns, 1)))\n\nprint(\"As noted above, recreation of input does not mean that recreation will be classified the same way. \\n\",\n      \"Using the above as input, the NN categorizes this as bird number \", x.argmax())","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.690805Z","iopub.status.idle":"2023-05-02T22:04:19.691317Z","shell.execute_reply.started":"2023-05-02T22:04:19.691078Z","shell.execute_reply":"2023-05-02T22:04:19.691108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", conv_data.min(), conv_data.max())\nuntrans_1 = Domain_to_Audio(conv_data, sr)\nprint(\"This is the recreated audio, (un-transformed) hear and see as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)\nIPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.694249Z","iopub.status.idle":"2023-05-02T22:04:19.69477Z","shell.execute_reply.started":"2023-05-02T22:04:19.694532Z","shell.execute_reply":"2023-05-02T22:04:19.69456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above is what the AI \"thinks bird class 1 - abhori1 - sounds like.\"\n* Many of the same observations regarding retained vversus discarded information","metadata":{}},{"cell_type":"markdown","source":"## Finally bird class 2 - \"abythr1\"\n* We'll hear what an example original audio sounds like\n* After that, we'll work backwards through the NN to approximate an input feature set (image) and from that audio best approximation\n* Listen to what the NN has learned each bird \"sounds like\" - you may be surprised!","metadata":{}},{"cell_type":"markdown","source":"### Listen to an example of what that bird sounds like","metadata":{}},{"cell_type":"code","source":"bird_class = 2\nexample_audio_index = 141\naudio_filename = image_data.at[example_audio_index,\"filename\"]\n# Image files are now in the \"input/v16e-gpu-all-birdclef2023-mindreader\" directory instead of \"working\"\naudio_filename = audio_filename.replace(\"working\", \"inputv16e-gpu-all-birdclef2023-mindreader\")\ncommon_name = image_data.at[example_audio_index,\"common_name\"]\n# load the audio data\naudio, sr = librosa.load(\"/kaggle/input/birdclef-2023/train_audio/\" + audio_filename, sr = sr)\n# Take first 5 second \"segment\"\naudio = audio[0:segment]\n\nprint(\"Here is an example of original audio of this \", common_name, \" bird sounds like\\n\")\nlibrosa.display.waveshow(audio, sr=sr)\n\nprint(\"This is the audio playback of the ORIGINAL audio file\")\nIPython.display.Audio(data = audio, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.696933Z","iopub.status.idle":"2023-05-02T22:04:19.697455Z","shell.execute_reply.started":"2023-05-02T22:04:19.697201Z","shell.execute_reply":"2023-05-02T22:04:19.697229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### For the above example, show the extracted features","metadata":{}},{"cell_type":"code","source":"# These are the features for the above audio sample\nfeat = Audio_to_Domain(audio, sr)\n# Move the range from current min and max, into 0 to 255 8 bit integers\nfmin = feat.min()\nfmax = feat.max()\nfrange = fmax - fmin\nfeat = np.array((((feat - fmin) / frange)*255), dtype='uint8')\nprint(\"These are teh features that will be used as input for the above example audio.\\n\")\nplt.imshow(feat, cmap='gray', aspect=int(num_columns/num_rows) , interpolation = 'None')\nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.699308Z","iopub.status.idle":"2023-05-02T22:04:19.699995Z","shell.execute_reply.started":"2023-05-02T22:04:19.699699Z","shell.execute_reply":"2023-05-02T22:04:19.699737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Force the output category of the NN to be the above bird","metadata":{}},{"cell_type":"code","source":"# Let's see if we can work backwards from a classification - in this case, we'll use the frst bird\nbird_class = 2 # You can pick a different bird here if you wish\nforce_class = np.zeros(num_birds)\nforce_class[bird_class] = 1\n\nprint(\"Here is the forced output of our classification layer (the last Dense layer with Softmax activation)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_birds)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, force_class)\n\n# Show plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.70202Z","iopub.status.idle":"2023-05-02T22:04:19.702572Z","shell.execute_reply.started":"2023-05-02T22:04:19.702298Z","shell.execute_reply":"2023-05-02T22:04:19.702324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's calculate the Embed layer for this forced classification\n# Weights from the embed layer to the Classification layer\nclass_weights = np.copy(model_class.get_layer(index=class_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", class_weights.shape)\n\nnum_embed_neurons = class_weights.shape[0]\nnum_class_neurons = class_weights.shape[1]\nembed_data = np.zeros(num_embed_neurons)\nfor i in range(num_embed_neurons):\n    for j in range(num_class_neurons):\n        embed_data[i] += force_class[j] * class_weights[i, j]\n\nprint(\"Here is the forced output of our embed layer (the first Dense layer after flattened last Conv layer)\")\n# Creating bar chart of Embed layer output values\nx = np.arange(num_embed_neurons)\n\nfig, ax = plt.subplots(figsize =(10, 2))\nax.bar(x, embed_data)\n\n# Show plot\nplt.show()\n\n# Weights from the embed layer to the Flatten-ed last Conv layer\nembed_weights = np.copy(model_embed.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the flattened last Conv layer to the embed has the shape \", embed_weights.shape)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.705093Z","iopub.status.idle":"2023-05-02T22:04:19.70561Z","shell.execute_reply.started":"2023-05-02T22:04:19.705352Z","shell.execute_reply":"2023-05-02T22:04:19.705377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Work backwards through the NN to recreate a best approximation of the input features, and audio, of the above bird.\n* Essentially \"What has the NN learned to be the audio feature set for the above bird, and what does it sound like?\"","metadata":{}},{"cell_type":"code","source":"# Let's calculate the last Conv2D layer for this forced classification\n# Weights from the last Conv2D layer to the embed layer\nlast_conv_weights_flat = np.copy(model.get_layer(index=embed_layer).get_weights()[0])\nprint(\"Weights array from the embed layer to the class layer has the shape \", last_conv_weights_flat.shape)\n\nlast_conv_weights = np.copy(np.reshape(last_conv_weights_flat, pre_flatten))\nprint(\"Weights array shape from the embed layer after undoing the Flatten \", last_conv_weights.shape)\n\n# The last convolutional network  is a 3 dimensional arrax of dimensions rows x columns x filters\nnum_conv_rows = last_conv_weights.shape[0]\nnum_conv_cols = last_conv_weights.shape[1]\nnum_conv_filt = last_conv_weights.shape[2]\n# In addition, each of these rows x columns x filter outputs is connected to a dense layer with a number of neurons (called the embed layer)\nnum_embed_neurons = last_conv_weights.shape[3]\n\n# As we build visualizations of each Convolutioal Layer, we cumulatively created \"filter patches\" that \n# represent all of the prior Convolutional layers as well. Because of this, we don't need to \n# recalculate each layer backwards. We can just use the forward calculated final convolutional layer\n# filter patch\nlast_conv_layer = num_conv - 1\nlast_filter_patches = filter_patches[last_conv_layer]\n\n# We will use the values of the embed layer (called embed_data here) multiplied by all\n# the corresponding weights that connected the last convolutional layer to that embed\n# layer (we took care of the \"Flatten\" above), and those values are the multiplier\n# we will apply to each filter patch at the respective location of the recreation.\n# In this way, we are moving backwards through the NN; from embed, to last Convolutional\n# layer, and finally all the way to a best approximation of the input.\n# NOTE: Since we used padding in the forward pass, we'll need to add a border to \n# our workspace input recreation.\nborder = int(last_filter_patches.shape[0]/2) \nextra_pixels = 2 * border \n\n# Create the workspace where individual filter patchess will be weighted and added up\nconv_data1 = np.zeros((num_conv_rows + extra_pixels, num_conv_cols + extra_pixels))\n# add 'em up\nfor i in range(num_embed_neurons):\n    embed_value = embed_data[i] # using local variables to speed things up\n    for j in range(num_conv_filt):\n        this_filter_patch = last_filter_patches[:,:,j]\n        this_conv_to_embed_weights = last_conv_weights[:,:,j,i]\n        for k in range(border, num_conv_cols + border, 1):\n            for l in range(border, num_conv_rows + border, 1):\n                conv_data1[l-border:l+border+1, k-border:k+border+1] += this_filter_patch[:,:] * this_conv_to_embed_weights[l-border,k-border] * embed_value\n                \nconv_data2 = np.copy(conv_data1)\nPIL_image = Image.fromarray(conv_data2)\nconv_data = np.array(PIL_image.resize([num_columns, num_rows]))\n\nprint(\"Here is the best approximation of the input for this bird\\n\")\nplt.imshow(conv_data, cmap=\"gray\", aspect=(num_columns/num_rows))\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.707313Z","iopub.status.idle":"2023-05-02T22:04:19.707878Z","shell.execute_reply.started":"2023-05-02T22:04:19.707622Z","shell.execute_reply":"2023-05-02T22:04:19.707648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = model.predict(np.reshape(conv_data, (1, num_rows, num_columns, 1)))\n\nprint(\"As noted above, recreation of input does not mean that recreation will be classified the same way. \\n\",\n      \"In fact, using the above as input, the NN categorizes this as bird number \", x.argmax())\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.710054Z","iopub.status.idle":"2023-05-02T22:04:19.710591Z","shell.execute_reply.started":"2023-05-02T22:04:19.710315Z","shell.execute_reply":"2023-05-02T22:04:19.71034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"After scaling, normalized transform has values within this range: \", conv_data.min(), conv_data.max())\nuntrans_1 = Domain_to_Audio(conv_data, sr)\nprint(\"This is the recreated audio, (un-transformed) hear and see as a graphic file\")\nlibrosa.display.waveshow(untrans_1, sr=sr)\nIPython.display.Audio(data = untrans_1, rate=sr)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.712258Z","iopub.status.idle":"2023-05-02T22:04:19.712877Z","shell.execute_reply.started":"2023-05-02T22:04:19.712527Z","shell.execute_reply":"2023-05-02T22:04:19.712553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The above is what the AI \"thinks bird class 2 - abythr1 - sounds like.\"\n* Many of the same observations regarding retained vversus discarded information","metadata":{}},{"cell_type":"markdown","source":"# Submit the Results\n* You will likely run this notebook twice or three times - first to create all the training set images, then to train a NN, and finally to use that trained NN to infer solutions to the contest.\n* I have included all of the code - including submission to the contest - to make this easier for you.\n\nBelow code modified from \"Inferring Birds with Kaggle Models\", courtesy PHIL CULLITON +1 · COPIED FROM PRIVATE NOTEBOOK +42,-5 · 1MO AGO · 22,483 VIEWS","metadata":{}},{"cell_type":"code","source":"def predict_for_sample(filename, sample_submission, frame_limit_secs=None):\n    file_id = filename.split(\".ogg\")[0].split(\"/\")[-1]\n    \n    audio, sample_rate = librosa.load(filename, sr = sr)\n        \n    # Get number of samples for 5 seconds\n    buffer = segment\n\n    samples_total = len(audio)\n    samples_segmented = 0\n    counter = 1\n\n    frame = 5\n\n    # this will be a list of 5 second segments\n    feature_split = []\n    classifications = []\n\n    if num_birds < len(competition_class_map):\n        print(\"Have to append 0% likelihood bird types because only \", num_birds, \"were used in taining\\n\")\n\n    while samples_segmented < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_segmented):\n            buffer = samples_total - samples_segmented\n\n        block = audio[samples_segmented : (samples_segmented + buffer)]\n\n        # convert the features \n        these_features = Audio_to_Domain(block, sr)\n\n        #scale to 0 - 255\n        fmax = these_features.max()\n        fmin = these_features.min()\n        frange = fmax - fmin\n            \n        these_features = np.array((((these_features - fmin) / frange)*255))\n        feature_split.append(these_features)\n\n        # Classfy each 5 secons segment\n        scratch = np.copy(np.reshape(these_features, (1,num_rows, num_columns, 1)))\n        scratch2 = (model.predict(scratch, verbose=0))\n        classifications.append(scratch2)\n        \n        counter += 1\n        samples_segmented += buffer\n                \n        probabilities = tf.nn.softmax(scratch2).numpy()\n        probabilities = np.copy(np.reshape(probabilities, (num_birds)))\n        if num_birds < len(competition_class_map) :\n            probabilities = np.append(probabilities, np.zeros(len(competition_class_map) - num_birds))\n            print(\"Had to append bird types because only \", num_birds, \"were used in taining\")\n        print(samples_segmented, samples_total)\n        \n        ## set the appropriate row in the sample submission\n        sample_submission.loc[sample_submission.row_id == file_id + \"_\" + str(frame), competition_classes] = probabilities[competition_class_map]\n        frame += 5\n        \ndef predict_for_sample(filename, sample_submission, frame_limit_secs=None):\n    file_id = filename.split(\".ogg\")[0].split(\"/\")[-1]\n    \n    audio, sample_rate = librosa.load(filename, sr = sr)\n        \n    # Get number of samples for 5 seconds\n    buffer = segment\n\n    samples_total = len(audio)\n    samples_segmented = 0\n    counter = 1\n\n    frame = 5\n\n    # this will be a list of 5 second segments\n    feature_split = []\n    classifications = []\n\n    if num_birds < len(competition_class_map):\n        print(\"Had to append bird types because only \", num_birds, \"were used in taining\")\n    while samples_segmented < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_segmented):\n            buffer = samples_total - samples_segmented\n\n        block = audio[samples_segmented : (samples_segmented + buffer)]\n\n        # convert the features \n        these_features = Audio_to_Domain(block, sr)\n\n        #scale to 0 - 255\n        fmax = these_features.max()\n        fmin = these_features.min()\n        frange = fmax - fmin\n            \n        these_features = np.array((((these_features - fmin) / frange)*255))\n        feature_split.append(these_features)\n\n        # Classfy each 5 second segment\n        scratch = np.copy(np.reshape(these_features, (1, num_rows, num_columns, 1)))\n        scratch2 = (model.predict(scratch, verbose = 1))\n        classifications.append(scratch2)\n        \n        counter += 1\n        samples_segmented += buffer\n                \n        probabilities = tf.nn.softmax(scratch2).numpy()\n        probabilities = np.copy(np.reshape(probabilities, (num_birds)))\n        if num_birds < len(competition_class_map) :\n            probabilities = np.append(probabilities, np.zeros(len(competition_class_map) - num_birds))\n        # print(samples_segmented, samples_total)\n        \n        ## set the appropriate row in the sample submission\n        sample_submission.loc[sample_submission.row_id == file_id + \"_\" + str(frame), competition_classes] = probabilities[competition_class_map]\n        frame += 5","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.720614Z","iopub.status.idle":"2023-05-02T22:04:19.72118Z","shell.execute_reply.started":"2023-05-02T22:04:19.720923Z","shell.execute_reply":"2023-05-02T22:04:19.720952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\ntest_samples = list(glob.glob(\"/kaggle/input/birdclef-2023/test_soundscapes/*.ogg\"))\ntest_samples","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.723739Z","iopub.status.idle":"2023-05-02T22:04:19.724251Z","shell.execute_reply.started":"2023-05-02T22:04:19.724015Z","shell.execute_reply":"2023-05-02T22:04:19.724042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_metadata = pd.read_csv(\"/kaggle/input/birdclef-2023/train_metadata.csv\")\ncompetition_classes = sorted(train_metadata.primary_label.unique())\n\ncompetition_class_map = []\nfor c in competition_classes:\n        i = competition_classes.index(c)\n        competition_class_map.append(i)\n        \n\nsample_sub = pd.read_csv(\"/kaggle/input/birdclef-2023/sample_submission.csv\")\nsample_sub[competition_classes] = sample_sub[competition_classes].astype(np.float32)\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.725603Z","iopub.status.idle":"2023-05-02T22:04:19.726081Z","shell.execute_reply.started":"2023-05-02T22:04:19.725858Z","shell.execute_reply":"2023-05-02T22:04:19.725883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frame_limit_secs = 15 if sample_sub.shape[0] == 3 else None\nfor sample_filename in test_samples:\n    predict_for_sample(sample_filename, sample_sub, frame_limit_secs=15)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.72891Z","iopub.status.idle":"2023-05-02T22:04:19.729584Z","shell.execute_reply.started":"2023-05-02T22:04:19.729266Z","shell.execute_reply":"2023-05-02T22:04:19.7293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.731502Z","iopub.status.idle":"2023-05-02T22:04:19.732036Z","shell.execute_reply.started":"2023-05-02T22:04:19.731806Z","shell.execute_reply":"2023-05-02T22:04:19.731831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T22:04:19.734136Z","iopub.status.idle":"2023-05-02T22:04:19.734776Z","shell.execute_reply.started":"2023-05-02T22:04:19.734424Z","shell.execute_reply":"2023-05-02T22:04:19.734523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Additional Background Information\nAs you know, in this notebook, we are taking the first step to generate a submission to the [BirdClef2023 competition](https://www.kaggle.com/c/birdclef-2023).  The goal of the competition is to identify Eastern African bird species by sound. \n### \"Reading the Robot Mind\" is a system that allows non-programmers to examine performance of individual and combined stages of the classification process\nMost important for this system is that it can be accomplished by a subject matter expert who may not be familair with the inner workings of AI. Subject matter experts know how to listen to birds and identify them, so they need data presented as audio (and possibly spectral analysis, as I've seen this used extensively by experts when exmining data). Additional references and history of the reading the robot mind system is explained in detail below. \n#### Stages of an automated process to classify birds:\nThe stages of the desired classification system include:\n\n* A method to record audio (provided with the contest)\n* A method of segmenting audio data (somewhat provided with the contest - 5 second segments)\n* A method of transforming the audio segment into a sequence of feature vectors (feature extraction such as Mel Scale Frequency Spectrum and MFCC)\n* A method of converting sequence of features into an image for CNN processing (in this case, also quantizing to 8-bit unsigned values)\n* A method where several CNN (Conv2D) and data reduction (MaxPooling) layers process data in sequence, iteratively extracting salient aspects\n* A method of taking a multi-layer CNN and outputting an embedding (Flatten then Dense layer)\n* A method of taking the embedding and converting it to a bird classification (Dense with softmax activation)\n\nPrior to the \"Reading the Robot Mind\" system described herein; a subject matter expert will not be able to examine issues associated with individual stages (such as listening to audio and looking at visualizations) except for the first two stages. Issues that can arise include: Inability to observe the difference between samples from different bird classifications. Inability to observe similarities between samples from the same bird classification. Inability to use audio and visual clues to correctly classify the bird.\n\nThis notebook presents a method whereby at any stage of the classification system, the expert can examine the internal workings of the system, and work backwards to audio and visual data in a way that is interpretable by a human expert.\n### We will qualitatively examine feature extraction algorithms as part of an effort to \"Read the Robot Mind\"\nThis notebook creates a method to allow users to select different hyper-paramenters (methods) at each stage and qualitatively compare them. Variations of segmentation algorithms are not presented in this notebook (since they are dictated to some extent by the rules of this competition).\n### Also examined is the Neural Network we trained. Can it too be used to \"Read the Robot Mind?\"\nHere we train a simple NN and see that it can classify with some level of accuracy. We can then use that NN \"in reverse\" to attempt to recreate an audio signal that generated a particular stage (often referred to as a layer) in the classification.\n### Finally - what if we simply supply a classification?\nThis notebook demonstrates the ability to simply provide the output - and then work backwards through the AI layers and reproduce a best approximation of the input that cause this classification. Effectively asking the AI system to \"Let me hear what you've learned a particular bird sounds like.\"","metadata":{}},{"cell_type":"markdown","source":"# Brief History of the \"Reading the Robot Mind\" System\nIt used to be that only human experts examined data and made decisions. Now Artificial Intelligence (AI) is enabling robotic decision making in an ever-widening variety of applications. As society allows this to happen, there is a greater likelihood that these robot decisions can affect people’s lives. It makes sense, therefore, to understand the capabilities and societal implications of AI robots.\n\nBig data is a term used to describe both the opportunities and the problems associated with so much information now available for decision making. With the advent of the Internet of Things (IoT) the impact of this huge amount of data is only growing. Actionable decisions need to be distilled from big data and AI can only go so far based on statistical analysis. Because of this, many deep learning algorithms are being developed. \n\n##  What exactly have these AI robots learned so deeply from all of this big data? \n\nThis question is very reasonable for society to ask. It is not enough to train and create a great AI robot. Many researchers are realizing that before their systems can be deployed, they must be able to prove to human experts that the robots learned the right things from the right data. This is difficult because human expert decision makers are not necessarily the same people who are good at creating robot AI. These two teams must work together in a user friendly way.\n## If only we could read the robot mind.\nRecently some old research of mine has been getting increased attention by engineers and scientists working on exactly the issues discussed above. This research yielded four US patents and a Windows app that gives a friendly user interface to allow examination of big data and inner workings of a trained AI decision maker. I never publically released the software, called INTEGRAT, because I sold rights to the patents to my employer at the time. \n\n## Now that the patents have expired, I can share my work.\n\n* For even more details, the reader is directed to the specific patent and independent claims listed below, all publically available at www.USPTO.gov.\n","metadata":{}},{"cell_type":"markdown","source":"# US Patent 5809462\nPaul A. Nussbaum, Inventor\n## Question: Is the feature extraction algorithm good and can the features be used as a codec?\n(\"codec\" is a portmanteau or combination of the words \"coder\" and \"decoder\", such as jpeg or mp3, etc...)\n\nAll AI systems train on data from the real world gathered by sensors, but before the data can make it to the AI, it must be digitized and important features must be extracted (sometimes called data transformation). \"Reading the Robot Mind\" allows a human expert to determine if the features extracted have retained sufficient information to identify the correct classification. \n\n* The question is: does the feature extraction algorithm retain classification-necessary information (or have we thrown away too much information)?\n\n* The premise is as follows: Human experts can observe the original signal and/or measurements and visualizations thereof - and correctly make the classification. \n\n* The thesis is that: A sufficient condition that the feature extraction algorithm isn't throwing away too much info as it seeks to condition and compress the signal; is that there exists a reverse transformation able to use the extracted features to recreate the original signal and/or measurements and visualizations thereof - such that the human expert will use these to correctly make the classification.\n\n* A bonus is: The human expert can also be presented the extracted features, and may find that they can be useful in future human classification exercises (observedly similar for same classifications, while different from others).\n\nSalient points from US Patent 5809462:\n\n* Independent Claim 1 – Provide a friendly user interface to present the features to the eyes and ears of the human expert to see if they can make the correct decisions based on the transformation features alone.\n   \n* Independent Claim 14 – Use that same user interface to let the human expert make sure that the data used to train one possible decision all seem similar to each other, and different from data which should yield different decisions.\n   \n* Independent Claim 25 – Let the human expert see if the features contain enough information to be used as a codec. In other words, if we can transform the original data into a set of features, we should be able to create original data (somewhat distorted and simplified perhaps) from the features – with sufficient information for a human expert to use that to identify the correct decision. \n","metadata":{}},{"cell_type":"markdown","source":"# US Patent 5867816\nPaul A. Nussbaum, Inventor\n## Question: Will the robot function just as well when the data segmentation, feature extraction, and decision identifier are fully automated?\n\nThe AI Robot is not useful if it requires a human expert to accompany the device when it is deployed in the field for day to day use. The robot must work on its own. Nevertheless, a quality assurance mechanism is required that makes the robot demonstrate it is working in “fully automatic” mode while a human expert grades it on the decisions the AI has made.\n\n* \"Reading the Robot Mind\" means that this segmentation, feature extraction, and classifier pipeline (inference) are observable in situ (on-site in real deployed situations, preferably live).\n\n* Bonus: The human expert can make adjustments to this pipeline to help get the desired classification accuracy, without requiring technical knowledge of the pipeline.\n\nIn the BirdClef2023 \n\nThe salient point from US Patent 5867816 is:\n\n* Independent Claims 1, 11, and 29 – The human expert examines and modifies the segmentation, features, and decision identifications to improve automated functioning.\n","metadata":{}},{"cell_type":"markdown","source":"# US Patent 5749066\nPaul A. Nussbaum, Inventor\n## Question: Is the training data good and can the classifier be used as a codec? \n\nAI robots can be finicky about the actual examples selected to be used for training. Giving too many of one kind of example can have a negative impact on the resulting decision maker, as can the introduction of misleading data examples (either accidentally or maliciously). Sometimes, a training set can be obviously bad, and other times, only after the AI is trained, can the human expert see that something has gone wrong.\n   \n* Independent Claims 1 and 7 - Provide a friendly user interface to present all stages of creating the training set to the eyes and ears of the human expert to see if they can spot any problems noted above.\n\n* Independent Claim 3 – Let the human expert see if the AI can be used as a codec. In other words, if the AI can identify the correct decision, then we can work backwards through the AI layers to recreate an exemplary feature set. We can then create the original data (somewhat distorted and simplified perhaps) from those features – with sufficient information for a human expert to use that to identify the correct decision.\n","metadata":{}},{"cell_type":"markdown","source":"#  US Patent 5864803\nThis Notebook DOES NOT cover these aspects in detail\n## Question: How can we assign more than one correct classification when experts cannot agree?\nBecause data sets can be very large and continuously streaming in from many sources, AI robot creators will chop up the data into segments that can be presented to identify a decision. Sometimes different human experts will see the same data segment and identify different decisions. Similarly, sometimes a human expert will examine one data segment and come up with two possible decisions, each of which are valid. INTEGRAT accounts for this situation in all stages of the creation and testing of the data and the AI robot.\n* Independent Claims 1 and 6 - Provide a friendly user interface to present data segments to the human expert and let them change the segmentation algorithm, or manually change a segment, if it could yield more than one decision.\n* Independent Claims 13, 17, and 20 – Provide a mechanism whereby multiple “correct” decisions can be assigned to one data segment, as well as AI training algorithms to support this ambiguity. Finally, allow the AI robot to come up with a single “best” decision when presented with ambiguous data, while also alerting users to “second best” decisions, and so on.\n","metadata":{}}]}