{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This is my attempt at visualizing the bird sounds. Send these spectrograms through your trusted CNN of choice coupled with a branch for the metadata and you should have a good starting point for a great result in this competition.\n\nThis version solves the performance issues I encountered before, so threading has been removed :-)\n\nOk. First the libraries:","metadata":{}},{"cell_type":"code","source":"from pydub import AudioSegment\nimport matplotlib.pyplot as plt\nfrom scipy.io import wavfile\nfrom tempfile import mkstemp\nimport numpy as np\nfrom pydub.utils import make_chunks\nimport glob, os\nfrom PIL import Image","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next the function that actually creates the spectrograms. This is what is called in a separate thread (one at a time in this implementation).\n\nFirst we read the .ogg file, then we break it up in 5 second chunks in the loop.\nI have implemented a check that we don't want to save more than 5 spectrograms from each file to speed it up a bit.\nLastly we use pyplot to render and save the spectrogram.","metadata":{}},{"cell_type":"code","source":"def create_spectrograms(filename):\n    print('./' + filename)\n    mp3_audio = AudioSegment.from_ogg('/' + filename)\n\n    sample = 0\n\n    chunk_length_ms = 5000\n    chunks = make_chunks(mp3_audio, chunk_length_ms)\n\n    for i, chunk in enumerate(chunks):\n        sample = sample + 1\n        if sample > 5:\n            sample = 0\n            break\n        fd, wname = mkstemp('.wav')\n        os.unlink(wname)\n        chunk.export(wname, format=\"wav\")\n        FS, data = wavfile.read(wname)\n        os.close(fd)\n        \n        data_size = np.array(data.shape).shape[0]\n        if data_size > 1:\n            plt.specgram(data[:, 0], Fs=FS, NFFT=128, noverlap=0)\n        else:\n            plt.specgram(data, Fs=FS, NFFT=128, noverlap=0)\n\n        plt.axis('off')\n        plt.savefig('spectrograms/' + fnames[5] + '/' + fnames[6].replace('.ogg', '_') + str(i * 5 + 5) + '.png',\n                    bbox_inches='tight')\n        plt.close()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok, now the final part which is the main part of the logic. Here we loop through the files one by one. If we encounter a new directory we create it in the spectrograms folder as well for easy training later. Otherwise it is a simpple matter of creating a thread with create_spectrograms, passing the file name we've reached.","metadata":{}},{"cell_type":"code","source":"failed = 0\ncount = 0\nold_path = ''\n\ntry:\n    os.mkdir('spectrograms/')\nexcept OSError:\n    print('Could not create spectrograms directory')\n\nfor filename in glob.iglob('/kaggle/input/birdclef-2021/train_short_audio/**', recursive=True):\n    if os.path.isfile(filename):\n        fnames = filename.split('/')\n\n        if old_path != fnames[5]:\n            # THIS IS ONLY TO ENSURE WE JUST LOOK AT THE FIRST FIVE SPECIES. \n            # YOU DON'T WANT THAT IN YOUR CODE FROM HERE...\n            count += 1\n            if count > 5:\n                break\n            old_path = fnames[5]\n            # TO HERE\n                \n            try:\n                os.mkdir('spectrograms/' + fnames[5])\n            except OSError:\n                failed = failed + 1\n        else:   # THIS IS ONLY TO ENSURE WE JUST LOOK AT THE FIRST SAMPLE OF EACH SPECIES. \n                # YOU DON'T WANT THAT IN YOUR CODE. FROM HERE...\n            continue\n            # TO HERE\n\n        create_spectrograms(filename)\n\nprint('All done...')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So now we can see the intricate details of the bird song (and train our network in recognizing these details).\n\nCompare for example the call of Bonaparte's Gull:","metadata":{}},{"cell_type":"code","source":"plt.imshow(Image.open('spectrograms/bongul/XC516224_20.png'))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With that of the Ruby-throated Hummingbird:","metadata":{}},{"cell_type":"code","source":"plt.imshow(Image.open('spectrograms/rthhum/XC319184_10.png'))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}