{
  "id": 451041,
  "title": "763rd Place Solution for the BirdCLEF 2023 Competition",
  "url": "/competitions/birdclef-2023/discussion/451041",
  "author_name": "Paul Nussbaum, PhD",
  "post_date": "2023-10-26T23:45:47.583000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>Context</strong></p>\n<ul>\n<li>Business context (use case) - <a href=\"https://www.kaggle.com/competitions/birdclef-2023/overview\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/overview</a></li>\n<li>Data context - <a href=\"https://www.kaggle.com/competitions/birdclef-2023/data\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/data</a><br>\n<strong>Overview of the Approach</strong><br>\nI have named this approach the Reading the Robot Mind system. Specifically, an interactive Jupyter notebook [1] is built for the purpose of training and deploying a deep learning neural network artificial intelligence (AI) [2] to automatically identify birds from recorded audio [3]. The notebook allows the user to modify parameters along the training and classification (inference) pipeline and observe the results. As with traditional observation methods, the notebook lets users view visual representations (spectrograms, etc.) of input vectors for similar and different birds [4]. In addition to traditional methods, this notebook also presents data in its original format (audio recordings of birds). This is common practice for a field researcher or subject matter expert (SME) testing a microphone and recording system [5]- they will want to listen to the recordings to see if they contain valid and sufficient information. The notebook [6] extends this intuitive and useful technique to individual neural network layers - working backwards towards a best estimate of the original input (referred to in this working note as \"reading the robot mind\"). The user can even provide just the \"answer\" (select a bird at the final output layer), and the reading the robot mind system will work backwards through the entire automated process and AI layers to let the SME hear a best approximation of what the AI has learned that bird sounds like.<br>\n<strong>Details of the Submission</strong><br>\nNote that due to the time and compute limitations imposed by the computing environment and contest rules provided, the notebook is divided into four public notebooks:</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader</a> - This notebook focuses on the Segmentation and Feature Extraction aspects of the AI solution, allowing users to make modifications and see and hear how much information is retained.</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v15h-all-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v15h-all-birdclef2023-mindreader</a> - This notebook allows the user to use their final decision related to segmentation and feature extraction, and convert and save all the BirdClef2023 data into this format.</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader</a> - This notebook uses the final decisions noted above, and trains the entire AI for a longer period of time, achieving better accuracy, and saving the trained AI system.</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v17b-all-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v17b-all-birdclef2023-mindreader</a> - This notebook brings all of this together for the sake of the contest submission and scoring.</li>\n</ul>\n<p>In the following sections, segmentation, feature extraction, quantization, AI model creation, training, and validation are discussed in detail. Also shown with formulas and examples are the aspects of the reading the robot mind system, including visualization of filters, recreation of input approximations based on outputs of intermediate and final layers of the AI system, and also the method whereby the output can be forced to an individual bird, and a best approximation of what that bird sounds like is created by the system.</p>\n<p><em>Segmentation, Feature Extraction, and Image Quantization Analysis</em><br>\nAn automated segmentation algorithm is used, however, due to the rules of the BirdCLEF 2023 competition, the SME is not permitted to modify this algorithm [3]. After segmentation, the audio data is transformed into another domain (called feature extraction in this document), and finally saved as a two-dimensional grayscale image with 8-bit quantization. The quantization was chosen due to the power and memory constraints of the edge device that will be performing inference (bird classification) in the field.<br>\nThe notebook allows the SME to try several Feature Extraction Algorithms and test the following:</p>\n<ul>\n<li>Do the features visually look similar for the same bird, and different for different birds?</li>\n<li>Is the similarity/difference enough to be able to visually classify which bird is which?</li>\n<li>If the feature extraction algorithm is performed in reverse to recreate the inputted audio (or an approximation thereof, due to the lossy nature of feature extraction) - is the recreated sound clear enough for the SME to identify the bird?</li>\n<li>If the answer is \"no\" to any of the above, allow fine tuning by the user</li>\n</ul>\n<p><em>Layer-wise Analysis of the AI (Convolutional, Max Pooling, and Dense Layers)</em><br>\nA simple sequential (not residual) convolutional AI is used, with dimensions and architecture shown in the aforementioned notebooks. This was found to yield a satisfactory accuracy measure for the purposes of demonstrating the reading the robot mind system, while still falling within the running time and memory constraints imposed by the development platform and contest rules.</p>\n<p><em>Visualizing Filter Patches</em><br>\nThe notebook allows visualization of the convolutional filter patches. Although this information is more useful to the AI programmer than it is to the SME, it is presented here since it is a step in the system of reading the robot mind. The algorithms used are similar to the \"expansion\" method [12].</p>\n<p><em>Using Filter Patches to Reconstruct Approximation of Input (for Convolutional and Max Pooling layers)</em><br>\nUsing convolutional neural networks to extract patterns from image data necessarily involves loss of information. Each convolutional layer has only a limited number of filters. Even though that limited number of filters is greater than would be needed to form an ortho-normal basis set (each filter having all 0 weights except for a weight of one at a different individual pixel), the AI is not being trained as an auto-encoder, and is instead being trained to minimize the loss when classifying birds. Also, information is being lost through the use of the \"relu\" activation function, which makes all negative output values a 0. Most of all, information is being lost at the max pooling layers.</p>\n<p><em>Working Backwards from a Bird Classification (What Does This Bird Sound Like?)</em><br>\nThe reading the robot mind system also allows one more function that can prove useful to the SME who is helping the AI programmer improve the system. This function is the ability to specify a particular output (bird classification) and work backwards through the entire AI to recreate an approximation of the original audio input, even when no input is provided.</p>\n<p><em>Conclusions</em><br>\nThe reading the robot mind system is implemented, allowing the SME to observe and qualitatively analyze the internal data flow of deep learning neural networks in a format familiar to them.<br>\nWith this, the SME can understand where in the pipeline the most information is being discarded by the AI system, and possibly help the programmer make improvements in future systems.</p>\n<p><strong>Sources</strong><br>\n[1]<br>\nJ. Perkel, \"Why Jupyter is data scientists' computational notebook of choice,\" Nature, vol. 563.7732, no. (2018), pp. 145-147, 2018.<br>\n[2]<br>\nIBM, \"AI vs. Machine Learning vs. Deep Learning v. Neural Networks: What's the Difference?,\" 2023. [Online]. Available: <a href=\"https://www.ibm.com/cloud/blog/ai-vs-machine-learning-vs-deep-learning-vs-neural-networks\" target=\"_blank\">https://www.ibm.com/cloud/blog/ai-vs-machine-learning-vs-deep-learning-vs-neural-networks</a>.<br>\n[3]<br>\nS. D. T. K. H. R. H. C. F. G. H. G. H. V. W. P. R. J. A. Kahl, \"Overview of BirdCLEF 2023: Automated bird species identification in Eastern Africa.,\" Working Notes of CLEF 2023 – Conference and Labs of the Evaluation Forum, 2023.<br>\n[4]<br>\nAudibon Society, \"Start using Spectrograms to Read Bird Songs and Calls,\" 2023. [Online]. Available: <a href=\"https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls\" target=\"_blank\">https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls</a>.<br>\n[5]<br>\nAcoustic Nature, \"Best way to record birdsong: Gear guide and tips,\" 2020. [Online]. Available: <a href=\"https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls\" target=\"_blank\">https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls</a>. [Accessed 2023].<br>\n[6]<br>\nP. Nussbaum, \"v15h BirdClef2023 Mindreader,\" 2023. [Online]. Available: <a href=\"https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader</a>.<br>\n[7]<br>\nC. B. L. P. S. K. H. G. B. D. D. M. J. E. C. L. T. L. R. C. M. Š. M. H. M. S. H. G. R. P. W.-P. V. H. K. T. D. I. E. P. B. H. M. A. Joly, \"Overview of LifeCLEF 2023: evaluation of ai models for the identification and prediction of birds, plants, snakes and fungi,\" in International Conference of the Cross-Language Evaluation Forum for European Languages, 2023.<br>\n[8]<br>\nJ. M. K. M. S. a. M. P. W. Wu, \"Explainable AI for Early Detection of Health Changes Via Streaming Clustering,\" in 2022 IEEE International Conference on Fuzzy Systems, Padua, 2022.<br>\n[9]<br>\nS. P. a. S. A. S. Sutthithatip, \"(Explainable) Artificial Intelligence in Aerospace Safety-Critical Systems,\" in IEEE Aerospace Conference, Big Sky. MT, 2022.<br>\n[10]<br>\nLibrosa, \"librosa.feature.melspectrogram,\" 2023. [Online]. Available: <a href=\"https://librosa.org/doc/main/generated/librosa.feature.melspectrogram.html\" target=\"_blank\">https://librosa.org/doc/main/generated/librosa.feature.melspectrogram.html</a>. [Accessed 2023].<br>\n[11]<br>\nLibrosa, \"librosa.feature.mfcc,\" 2023. [Online]. Available: <a href=\"https://librosa.org/doc/main/generated/librosa.feature.mfcc.html\" target=\"_blank\">https://librosa.org/doc/main/generated/librosa.feature.mfcc.html</a>. [Accessed 2023].<br>\n[12]<br>\ne. a. Voss, \"Visualizing Weights,\" Distill, vol. 10.23915/distill.00024.007, no. 10.23915/distill.00024.007, p. 10.23915/distill.00024.007, 2021.<br>\n[13]<br>\nGoogle, 2023. [Online]. Available: <a href=\"https://distill.pub/2020/circuits/visualizing-weights/\" target=\"_blank\">https://distill.pub/2020/circuits/visualizing-weights/</a>.</p>",
  "messages": [
    {
      "id": 2500692,
      "postDate": "2023-10-26T23:45:47.583Z",
      "content": "<p><strong>Context</strong></p>\n<ul>\n<li>Business context (use case) - <a href=\"https://www.kaggle.com/competitions/birdclef-2023/overview\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/overview</a></li>\n<li>Data context - <a href=\"https://www.kaggle.com/competitions/birdclef-2023/data\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/data</a><br>\n<strong>Overview of the Approach</strong><br>\nI have named this approach the Reading the Robot Mind system. Specifically, an interactive Jupyter notebook [1] is built for the purpose of training and deploying a deep learning neural network artificial intelligence (AI) [2] to automatically identify birds from recorded audio [3]. The notebook allows the user to modify parameters along the training and classification (inference) pipeline and observe the results. As with traditional observation methods, the notebook lets users view visual representations (spectrograms, etc.) of input vectors for similar and different birds [4]. In addition to traditional methods, this notebook also presents data in its original format (audio recordings of birds). This is common practice for a field researcher or subject matter expert (SME) testing a microphone and recording system [5]- they will want to listen to the recordings to see if they contain valid and sufficient information. The notebook [6] extends this intuitive and useful technique to individual neural network layers - working backwards towards a best estimate of the original input (referred to in this working note as \"reading the robot mind\"). The user can even provide just the \"answer\" (select a bird at the final output layer), and the reading the robot mind system will work backwards through the entire automated process and AI layers to let the SME hear a best approximation of what the AI has learned that bird sounds like.<br>\n<strong>Details of the Submission</strong><br>\nNote that due to the time and compute limitations imposed by the computing environment and contest rules provided, the notebook is divided into four public notebooks:</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader</a> - This notebook focuses on the Segmentation and Feature Extraction aspects of the AI solution, allowing users to make modifications and see and hear how much information is retained.</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v15h-all-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v15h-all-birdclef2023-mindreader</a> - This notebook allows the user to use their final decision related to segmentation and feature extraction, and convert and save all the BirdClef2023 data into this format.</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader</a> - This notebook uses the final decisions noted above, and trains the entire AI for a longer period of time, achieving better accuracy, and saving the trained AI system.</li>\n<li><a href=\"https://www.kaggle.com/code/pnussbaum/v17b-all-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v17b-all-birdclef2023-mindreader</a> - This notebook brings all of this together for the sake of the contest submission and scoring.</li>\n</ul>\n<p>In the following sections, segmentation, feature extraction, quantization, AI model creation, training, and validation are discussed in detail. Also shown with formulas and examples are the aspects of the reading the robot mind system, including visualization of filters, recreation of input approximations based on outputs of intermediate and final layers of the AI system, and also the method whereby the output can be forced to an individual bird, and a best approximation of what that bird sounds like is created by the system.</p>\n<p><em>Segmentation, Feature Extraction, and Image Quantization Analysis</em><br>\nAn automated segmentation algorithm is used, however, due to the rules of the BirdCLEF 2023 competition, the SME is not permitted to modify this algorithm [3]. After segmentation, the audio data is transformed into another domain (called feature extraction in this document), and finally saved as a two-dimensional grayscale image with 8-bit quantization. The quantization was chosen due to the power and memory constraints of the edge device that will be performing inference (bird classification) in the field.<br>\nThe notebook allows the SME to try several Feature Extraction Algorithms and test the following:</p>\n<ul>\n<li>Do the features visually look similar for the same bird, and different for different birds?</li>\n<li>Is the similarity/difference enough to be able to visually classify which bird is which?</li>\n<li>If the feature extraction algorithm is performed in reverse to recreate the inputted audio (or an approximation thereof, due to the lossy nature of feature extraction) - is the recreated sound clear enough for the SME to identify the bird?</li>\n<li>If the answer is \"no\" to any of the above, allow fine tuning by the user</li>\n</ul>\n<p><em>Layer-wise Analysis of the AI (Convolutional, Max Pooling, and Dense Layers)</em><br>\nA simple sequential (not residual) convolutional AI is used, with dimensions and architecture shown in the aforementioned notebooks. This was found to yield a satisfactory accuracy measure for the purposes of demonstrating the reading the robot mind system, while still falling within the running time and memory constraints imposed by the development platform and contest rules.</p>\n<p><em>Visualizing Filter Patches</em><br>\nThe notebook allows visualization of the convolutional filter patches. Although this information is more useful to the AI programmer than it is to the SME, it is presented here since it is a step in the system of reading the robot mind. The algorithms used are similar to the \"expansion\" method [12].</p>\n<p><em>Using Filter Patches to Reconstruct Approximation of Input (for Convolutional and Max Pooling layers)</em><br>\nUsing convolutional neural networks to extract patterns from image data necessarily involves loss of information. Each convolutional layer has only a limited number of filters. Even though that limited number of filters is greater than would be needed to form an ortho-normal basis set (each filter having all 0 weights except for a weight of one at a different individual pixel), the AI is not being trained as an auto-encoder, and is instead being trained to minimize the loss when classifying birds. Also, information is being lost through the use of the \"relu\" activation function, which makes all negative output values a 0. Most of all, information is being lost at the max pooling layers.</p>\n<p><em>Working Backwards from a Bird Classification (What Does This Bird Sound Like?)</em><br>\nThe reading the robot mind system also allows one more function that can prove useful to the SME who is helping the AI programmer improve the system. This function is the ability to specify a particular output (bird classification) and work backwards through the entire AI to recreate an approximation of the original audio input, even when no input is provided.</p>\n<p><em>Conclusions</em><br>\nThe reading the robot mind system is implemented, allowing the SME to observe and qualitatively analyze the internal data flow of deep learning neural networks in a format familiar to them.<br>\nWith this, the SME can understand where in the pipeline the most information is being discarded by the AI system, and possibly help the programmer make improvements in future systems.</p>\n<p><strong>Sources</strong><br>\n[1]<br>\nJ. Perkel, \"Why Jupyter is data scientists' computational notebook of choice,\" Nature, vol. 563.7732, no. (2018), pp. 145-147, 2018.<br>\n[2]<br>\nIBM, \"AI vs. Machine Learning vs. Deep Learning v. Neural Networks: What's the Difference?,\" 2023. [Online]. Available: <a href=\"https://www.ibm.com/cloud/blog/ai-vs-machine-learning-vs-deep-learning-vs-neural-networks\" target=\"_blank\">https://www.ibm.com/cloud/blog/ai-vs-machine-learning-vs-deep-learning-vs-neural-networks</a>.<br>\n[3]<br>\nS. D. T. K. H. R. H. C. F. G. H. G. H. V. W. P. R. J. A. Kahl, \"Overview of BirdCLEF 2023: Automated bird species identification in Eastern Africa.,\" Working Notes of CLEF 2023 – Conference and Labs of the Evaluation Forum, 2023.<br>\n[4]<br>\nAudibon Society, \"Start using Spectrograms to Read Bird Songs and Calls,\" 2023. [Online]. Available: <a href=\"https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls\" target=\"_blank\">https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls</a>.<br>\n[5]<br>\nAcoustic Nature, \"Best way to record birdsong: Gear guide and tips,\" 2020. [Online]. Available: <a href=\"https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls\" target=\"_blank\">https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls</a>. [Accessed 2023].<br>\n[6]<br>\nP. Nussbaum, \"v15h BirdClef2023 Mindreader,\" 2023. [Online]. Available: <a href=\"https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader\" target=\"_blank\">https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader</a>.<br>\n[7]<br>\nC. B. L. P. S. K. H. G. B. D. D. M. J. E. C. L. T. L. R. C. M. Š. M. H. M. S. H. G. R. P. W.-P. V. H. K. T. D. I. E. P. B. H. M. A. Joly, \"Overview of LifeCLEF 2023: evaluation of ai models for the identification and prediction of birds, plants, snakes and fungi,\" in International Conference of the Cross-Language Evaluation Forum for European Languages, 2023.<br>\n[8]<br>\nJ. M. K. M. S. a. M. P. W. Wu, \"Explainable AI for Early Detection of Health Changes Via Streaming Clustering,\" in 2022 IEEE International Conference on Fuzzy Systems, Padua, 2022.<br>\n[9]<br>\nS. P. a. S. A. S. Sutthithatip, \"(Explainable) Artificial Intelligence in Aerospace Safety-Critical Systems,\" in IEEE Aerospace Conference, Big Sky. MT, 2022.<br>\n[10]<br>\nLibrosa, \"librosa.feature.melspectrogram,\" 2023. [Online]. Available: <a href=\"https://librosa.org/doc/main/generated/librosa.feature.melspectrogram.html\" target=\"_blank\">https://librosa.org/doc/main/generated/librosa.feature.melspectrogram.html</a>. [Accessed 2023].<br>\n[11]<br>\nLibrosa, \"librosa.feature.mfcc,\" 2023. [Online]. Available: <a href=\"https://librosa.org/doc/main/generated/librosa.feature.mfcc.html\" target=\"_blank\">https://librosa.org/doc/main/generated/librosa.feature.mfcc.html</a>. [Accessed 2023].<br>\n[12]<br>\ne. a. Voss, \"Visualizing Weights,\" Distill, vol. 10.23915/distill.00024.007, no. 10.23915/distill.00024.007, p. 10.23915/distill.00024.007, 2021.<br>\n[13]<br>\nGoogle, 2023. [Online]. Available: <a href=\"https://distill.pub/2020/circuits/visualizing-weights/\" target=\"_blank\">https://distill.pub/2020/circuits/visualizing-weights/</a>.</p>",
      "rawMarkdown": "**Context**\n- Business context (use case) - https://www.kaggle.com/competitions/birdclef-2023/overview\n- Data context - https://www.kaggle.com/competitions/birdclef-2023/data\n**Overview of the Approach**\nI have named this approach the Reading the Robot Mind system. Specifically, an interactive Jupyter notebook [1] is built for the purpose of training and deploying a deep learning neural network artificial intelligence (AI) [2] to automatically identify birds from recorded audio [3]. The notebook allows the user to modify parameters along the training and classification (inference) pipeline and observe the results. As with traditional observation methods, the notebook lets users view visual representations (spectrograms, etc.) of input vectors for similar and different birds [4]. In addition to traditional methods, this notebook also presents data in its original format (audio recordings of birds). This is common practice for a field researcher or subject matter expert (SME) testing a microphone and recording system [5]- they will want to listen to the recordings to see if they contain valid and sufficient information. The notebook [6] extends this intuitive and useful technique to individual neural network layers - working backwards towards a best estimate of the original input (referred to in this working note as \"reading the robot mind\"). The user can even provide just the \"answer\" (select a bird at the final output layer), and the reading the robot mind system will work backwards through the entire automated process and AI layers to let the SME hear a best approximation of what the AI has learned that bird sounds like.\n**Details of the Submission**\nNote that due to the time and compute limitations imposed by the computing environment and contest rules provided, the notebook is divided into four public notebooks:\n- https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader - This notebook focuses on the Segmentation and Feature Extraction aspects of the AI solution, allowing users to make modifications and see and hear how much information is retained.\n- https://www.kaggle.com/code/pnussbaum/v15h-all-birdclef2023-mindreader - This notebook allows the user to use their final decision related to segmentation and feature extraction, and convert and save all the BirdClef2023 data into this format.\n- https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader - This notebook uses the final decisions noted above, and trains the entire AI for a longer period of time, achieving better accuracy, and saving the trained AI system.\n- https://www.kaggle.com/code/pnussbaum/v17b-all-birdclef2023-mindreader - This notebook brings all of this together for the sake of the contest submission and scoring.\n\nIn the following sections, segmentation, feature extraction, quantization, AI model creation, training, and validation are discussed in detail. Also shown with formulas and examples are the aspects of the reading the robot mind system, including visualization of filters, recreation of input approximations based on outputs of intermediate and final layers of the AI system, and also the method whereby the output can be forced to an individual bird, and a best approximation of what that bird sounds like is created by the system.\n\n*Segmentation, Feature Extraction, and Image Quantization Analysis*\nAn automated segmentation algorithm is used, however, due to the rules of the BirdCLEF 2023 competition, the SME is not permitted to modify this algorithm [3]. After segmentation, the audio data is transformed into another domain (called feature extraction in this document), and finally saved as a two-dimensional grayscale image with 8-bit quantization. The quantization was chosen due to the power and memory constraints of the edge device that will be performing inference (bird classification) in the field.\nThe notebook allows the SME to try several Feature Extraction Algorithms and test the following:\n- Do the features visually look similar for the same bird, and different for different birds?\n- Is the similarity/difference enough to be able to visually classify which bird is which?\n- If the feature extraction algorithm is performed in reverse to recreate the inputted audio (or an approximation thereof, due to the lossy nature of feature extraction) - is the recreated sound clear enough for the SME to identify the bird?\n- If the answer is \"no\" to any of the above, allow fine tuning by the user\n\n*Layer-wise Analysis of the AI (Convolutional, Max Pooling, and Dense Layers)*\nA simple sequential (not residual) convolutional AI is used, with dimensions and architecture shown in the aforementioned notebooks. This was found to yield a satisfactory accuracy measure for the purposes of demonstrating the reading the robot mind system, while still falling within the running time and memory constraints imposed by the development platform and contest rules.\n\n*Visualizing Filter Patches*\nThe notebook allows visualization of the convolutional filter patches. Although this information is more useful to the AI programmer than it is to the SME, it is presented here since it is a step in the system of reading the robot mind. The algorithms used are similar to the \"expansion\" method [12].\n\n*Using Filter Patches to Reconstruct Approximation of Input (for Convolutional and Max Pooling layers)*\nUsing convolutional neural networks to extract patterns from image data necessarily involves loss of information. Each convolutional layer has only a limited number of filters. Even though that limited number of filters is greater than would be needed to form an ortho-normal basis set (each filter having all 0 weights except for a weight of one at a different individual pixel), the AI is not being trained as an auto-encoder, and is instead being trained to minimize the loss when classifying birds. Also, information is being lost through the use of the \"relu\" activation function, which makes all negative output values a 0. Most of all, information is being lost at the max pooling layers.\n\n*Working Backwards from a Bird Classification (What Does This Bird Sound Like?)*\nThe reading the robot mind system also allows one more function that can prove useful to the SME who is helping the AI programmer improve the system. This function is the ability to specify a particular output (bird classification) and work backwards through the entire AI to recreate an approximation of the original audio input, even when no input is provided.\n\n*Conclusions*\nThe reading the robot mind system is implemented, allowing the SME to observe and qualitatively analyze the internal data flow of deep learning neural networks in a format familiar to them.\nWith this, the SME can understand where in the pipeline the most information is being discarded by the AI system, and possibly help the programmer make improvements in future systems.\n\n**Sources**\n[1]\nJ. Perkel, \"Why Jupyter is data scientists' computational notebook of choice,\" Nature, vol. 563.7732, no. (2018), pp. 145-147, 2018.\n[2]\nIBM, \"AI vs. Machine Learning vs. Deep Learning v. Neural Networks: What's the Difference?,\" 2023. [Online]. Available: https://www.ibm.com/cloud/blog/ai-vs-machine-learning-vs-deep-learning-vs-neural-networks.\n[3]\nS. D. T. K. H. R. H. C. F. G. H. G. H. V. W. P. R. J. A. Kahl, \"Overview of BirdCLEF 2023: Automated bird species identification in Eastern Africa.,\" Working Notes of CLEF 2023 – Conference and Labs of the Evaluation Forum, 2023.\n[4]\nAudibon Society, \"Start using Spectrograms to Read Bird Songs and Calls,\" 2023. [Online]. Available: https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls.\n[5]\nAcoustic Nature, \"Best way to record birdsong: Gear guide and tips,\" 2020. [Online]. Available: https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls. [Accessed 2023].\n[6]\nP. Nussbaum, \"v15h BirdClef2023 Mindreader,\" 2023. [Online]. Available: https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader.\n[7]\nC. B. L. P. S. K. H. G. B. D. D. M. J. E. C. L. T. L. R. C. M. Š. M. H. M. S. H. G. R. P. W.-P. V. H. K. T. D. I. E. P. B. H. M. A. Joly, \"Overview of LifeCLEF 2023: evaluation of ai models for the identification and prediction of birds, plants, snakes and fungi,\" in International Conference of the Cross-Language Evaluation Forum for European Languages, 2023.\n[8]\nJ. M. K. M. S. a. M. P. W. Wu, \"Explainable AI for Early Detection of Health Changes Via Streaming Clustering,\" in 2022 IEEE International Conference on Fuzzy Systems, Padua, 2022.\n[9]\nS. P. a. S. A. S. Sutthithatip, \"(Explainable) Artificial Intelligence in Aerospace Safety-Critical Systems,\" in IEEE Aerospace Conference, Big Sky. MT, 2022.\n[10]\nLibrosa, \"librosa.feature.melspectrogram,\" 2023. [Online]. Available: https://librosa.org/doc/main/generated/librosa.feature.melspectrogram.html. [Accessed 2023].\n[11]\nLibrosa, \"librosa.feature.mfcc,\" 2023. [Online]. Available: https://librosa.org/doc/main/generated/librosa.feature.mfcc.html. [Accessed 2023].\n[12]\ne. a. Voss, \"Visualizing Weights,\" Distill, vol. 10.23915/distill.00024.007, no. 10.23915/distill.00024.007, p. 10.23915/distill.00024.007, 2021.\n[13]\nGoogle, 2023. [Online]. Available: https://distill.pub/2020/circuits/visualizing-weights/.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2500692": "**Context**\n- Business context (use case) - https://www.kaggle.com/competitions/birdclef-2023/overview\n- Data context - https://www.kaggle.com/competitions/birdclef-2023/data\n**Overview of the Approach**\nI have named this approach the Reading the Robot Mind system. Specifically, an interactive Jupyter notebook [1] is built for the purpose of training and deploying a deep learning neural network artificial intelligence (AI) [2] to automatically identify birds from recorded audio [3]. The notebook allows the user to modify parameters along the training and classification (inference) pipeline and observe the results. As with traditional observation methods, the notebook lets users view visual representations (spectrograms, etc.) of input vectors for similar and different birds [4]. In addition to traditional methods, this notebook also presents data in its original format (audio recordings of birds). This is common practice for a field researcher or subject matter expert (SME) testing a microphone and recording system [5]- they will want to listen to the recordings to see if they contain valid and sufficient information. The notebook [6] extends this intuitive and useful technique to individual neural network layers - working backwards towards a best estimate of the original input (referred to in this working note as \"reading the robot mind\"). The user can even provide just the \"answer\" (select a bird at the final output layer), and the reading the robot mind system will work backwards through the entire automated process and AI layers to let the SME hear a best approximation of what the AI has learned that bird sounds like.\n**Details of the Submission**\nNote that due to the time and compute limitations imposed by the computing environment and contest rules provided, the notebook is divided into four public notebooks:\n- https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader - This notebook focuses on the Segmentation and Feature Extraction aspects of the AI solution, allowing users to make modifications and see and hear how much information is retained.\n- https://www.kaggle.com/code/pnussbaum/v15h-all-birdclef2023-mindreader - This notebook allows the user to use their final decision related to segmentation and feature extraction, and convert and save all the BirdClef2023 data into this format.\n- https://www.kaggle.com/code/pnussbaum/v16e-gpu-all-birdclef2023-mindreader - This notebook uses the final decisions noted above, and trains the entire AI for a longer period of time, achieving better accuracy, and saving the trained AI system.\n- https://www.kaggle.com/code/pnussbaum/v17b-all-birdclef2023-mindreader - This notebook brings all of this together for the sake of the contest submission and scoring.\n\nIn the following sections, segmentation, feature extraction, quantization, AI model creation, training, and validation are discussed in detail. Also shown with formulas and examples are the aspects of the reading the robot mind system, including visualization of filters, recreation of input approximations based on outputs of intermediate and final layers of the AI system, and also the method whereby the output can be forced to an individual bird, and a best approximation of what that bird sounds like is created by the system.\n\n*Segmentation, Feature Extraction, and Image Quantization Analysis*\nAn automated segmentation algorithm is used, however, due to the rules of the BirdCLEF 2023 competition, the SME is not permitted to modify this algorithm [3]. After segmentation, the audio data is transformed into another domain (called feature extraction in this document), and finally saved as a two-dimensional grayscale image with 8-bit quantization. The quantization was chosen due to the power and memory constraints of the edge device that will be performing inference (bird classification) in the field.\nThe notebook allows the SME to try several Feature Extraction Algorithms and test the following:\n- Do the features visually look similar for the same bird, and different for different birds?\n- Is the similarity/difference enough to be able to visually classify which bird is which?\n- If the feature extraction algorithm is performed in reverse to recreate the inputted audio (or an approximation thereof, due to the lossy nature of feature extraction) - is the recreated sound clear enough for the SME to identify the bird?\n- If the answer is \"no\" to any of the above, allow fine tuning by the user\n\n*Layer-wise Analysis of the AI (Convolutional, Max Pooling, and Dense Layers)*\nA simple sequential (not residual) convolutional AI is used, with dimensions and architecture shown in the aforementioned notebooks. This was found to yield a satisfactory accuracy measure for the purposes of demonstrating the reading the robot mind system, while still falling within the running time and memory constraints imposed by the development platform and contest rules.\n\n*Visualizing Filter Patches*\nThe notebook allows visualization of the convolutional filter patches. Although this information is more useful to the AI programmer than it is to the SME, it is presented here since it is a step in the system of reading the robot mind. The algorithms used are similar to the \"expansion\" method [12].\n\n*Using Filter Patches to Reconstruct Approximation of Input (for Convolutional and Max Pooling layers)*\nUsing convolutional neural networks to extract patterns from image data necessarily involves loss of information. Each convolutional layer has only a limited number of filters. Even though that limited number of filters is greater than would be needed to form an ortho-normal basis set (each filter having all 0 weights except for a weight of one at a different individual pixel), the AI is not being trained as an auto-encoder, and is instead being trained to minimize the loss when classifying birds. Also, information is being lost through the use of the \"relu\" activation function, which makes all negative output values a 0. Most of all, information is being lost at the max pooling layers.\n\n*Working Backwards from a Bird Classification (What Does This Bird Sound Like?)*\nThe reading the robot mind system also allows one more function that can prove useful to the SME who is helping the AI programmer improve the system. This function is the ability to specify a particular output (bird classification) and work backwards through the entire AI to recreate an approximation of the original audio input, even when no input is provided.\n\n*Conclusions*\nThe reading the robot mind system is implemented, allowing the SME to observe and qualitatively analyze the internal data flow of deep learning neural networks in a format familiar to them.\nWith this, the SME can understand where in the pipeline the most information is being discarded by the AI system, and possibly help the programmer make improvements in future systems.\n\n**Sources**\n[1]\nJ. Perkel, \"Why Jupyter is data scientists' computational notebook of choice,\" Nature, vol. 563.7732, no. (2018), pp. 145-147, 2018.\n[2]\nIBM, \"AI vs. Machine Learning vs. Deep Learning v. Neural Networks: What's the Difference?,\" 2023. [Online]. Available: https://www.ibm.com/cloud/blog/ai-vs-machine-learning-vs-deep-learning-vs-neural-networks.\n[3]\nS. D. T. K. H. R. H. C. F. G. H. G. H. V. W. P. R. J. A. Kahl, \"Overview of BirdCLEF 2023: Automated bird species identification in Eastern Africa.,\" Working Notes of CLEF 2023 – Conference and Labs of the Evaluation Forum, 2023.\n[4]\nAudibon Society, \"Start using Spectrograms to Read Bird Songs and Calls,\" 2023. [Online]. Available: https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls.\n[5]\nAcoustic Nature, \"Best way to record birdsong: Gear guide and tips,\" 2020. [Online]. Available: https://www.audubon.org/news/start-using-spectrograms-read-bird-songs-and-calls. [Accessed 2023].\n[6]\nP. Nussbaum, \"v15h BirdClef2023 Mindreader,\" 2023. [Online]. Available: https://www.kaggle.com/code/pnussbaum/v15h-birdclef2023-mindreader.\n[7]\nC. B. L. P. S. K. H. G. B. D. D. M. J. E. C. L. T. L. R. C. M. Š. M. H. M. S. H. G. R. P. W.-P. V. H. K. T. D. I. E. P. B. H. M. A. Joly, \"Overview of LifeCLEF 2023: evaluation of ai models for the identification and prediction of birds, plants, snakes and fungi,\" in International Conference of the Cross-Language Evaluation Forum for European Languages, 2023.\n[8]\nJ. M. K. M. S. a. M. P. W. Wu, \"Explainable AI for Early Detection of Health Changes Via Streaming Clustering,\" in 2022 IEEE International Conference on Fuzzy Systems, Padua, 2022.\n[9]\nS. P. a. S. A. S. Sutthithatip, \"(Explainable) Artificial Intelligence in Aerospace Safety-Critical Systems,\" in IEEE Aerospace Conference, Big Sky. MT, 2022.\n[10]\nLibrosa, \"librosa.feature.melspectrogram,\" 2023. [Online]. Available: https://librosa.org/doc/main/generated/librosa.feature.melspectrogram.html. [Accessed 2023].\n[11]\nLibrosa, \"librosa.feature.mfcc,\" 2023. [Online]. Available: https://librosa.org/doc/main/generated/librosa.feature.mfcc.html. [Accessed 2023].\n[12]\ne. a. Voss, \"Visualizing Weights,\" Distill, vol. 10.23915/distill.00024.007, no. 10.23915/distill.00024.007, p. 10.23915/distill.00024.007, 2021.\n[13]\nGoogle, 2023. [Online]. Available: https://distill.pub/2020/circuits/visualizing-weights/."
  }
}