{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Making a sample submission\nIn this notebook, we will introduce the very basics of how to run inference on a hidden test set and submit detection results to the competition.\n\nFirst thing we need to do is to look for test soundscapes. **The hidden test set will only appear if you submit the notebook.** Yet, you can use one test soundscapes to validate your workflow.","metadata":{}},{"cell_type":"code","source":"import os\nimport json\nimport numpy as np\nimport pandas as pd\nimport librosa\n\n# First, load list of audio files by parsing the test_soundscape folder.\ntest_audio_dir = '../input/birdclef-2023/test_soundscapes/'\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\n# At the moment, there should only be a single soundscape visible.\n# During the submission re-run, all other hidden soundscapes\n# will be visible too and can be processed by your notebook.\nprint('Number of test soundscapes:', len(file_list))","metadata":{"execution":{"iopub.status.busy":"2023-03-06T11:54:02.000622Z","iopub.execute_input":"2023-03-06T11:54:02.001123Z","iopub.status.idle":"2023-03-06T11:54:02.050929Z","shell.execute_reply.started":"2023-03-06T11:54:02.001084Z","shell.execute_reply":"2023-03-06T11:54:02.049401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we need to open each test soundscape, process it, make predictions with our model, and store the results for all birds. In this notebook, we will not use a trained model. Instead, we will give dummy scores to each species and randomly assign a score for the submissions file.","metadata":{}},{"cell_type":"code","source":"# This is where we will store our results\npred = {'row_id': []}\ntrain_audio_dir = '../input/birdclef-2023/train_audio/'\nspecies_list = sorted(os.listdir(train_audio_dir))\nfor species_code in species_list:\n    pred[species_code] = []\n\n# Process audio files and make predictions\nfor afile in file_list:\n    \n    # Complete file path\n    path = test_audio_dir + afile + '.ogg'\n    \n    # Open file with librosa and split signal into 5-second chunks\n    # sig, rate = librosa.load(path, sr=32000)\n    # ...\n    \n    # Let's assume we have a list of 120 audio chunks (10min / 5s == 120 segments)\n    chunks = [[] for i in range(120)]\n    \n    # Make prediction for each chunk\n    # Each bird gets a random value in our case\n    # since we don't actually have a model\n    for i in range(len(chunks)):        \n        chunk_end_time = (i + 1) * 5\n        \n        # Assign the row_id which we need to do for each chunk\n        row_id = afile + '_' + str(chunk_end_time)\n        pred['row_id'].append(row_id)\n        \n        for bird in species_list:\n            \n            # This is our random prediction score for this bird\n            score = np.random.uniform()     \n            \n            # Put the result into our prediction dict            \n            pred[bird].append(score)","metadata":{"execution":{"iopub.status.busy":"2023-03-06T11:54:09.076150Z","iopub.execute_input":"2023-03-06T11:54:09.076660Z","iopub.status.idle":"2023-03-06T11:54:09.211322Z","shell.execute_reply.started":"2023-03-06T11:54:09.076615Z","shell.execute_reply":"2023-03-06T11:54:09.209986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, we need to save our results to a csv-file named 'submission.csv'.\n\n**Important:** Make sure to include a score for *all* birds for *every* 5-second segment of *every* file. If the number of rows and columns in your 'submission.csv' doesn't match the ground truth, the submission will fail.","metadata":{}},{"cell_type":"code","source":"# Make a new data frame and look at some results        \nresults = pd.DataFrame(pred, columns = ['row_id'] + species_list)\n\n# Quick sanity check\nprint(results.head()) \n    \n# Convert our results to csv\nresults.to_csv(\"submission.csv\", index=False)    ","metadata":{"execution":{"iopub.status.busy":"2023-03-06T11:54:12.470734Z","iopub.execute_input":"2023-03-06T11:54:12.471380Z","iopub.status.idle":"2023-03-06T11:54:12.594572Z","shell.execute_reply.started":"2023-03-06T11:54:12.471325Z","shell.execute_reply":"2023-03-06T11:54:12.593425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we are ready to submit, and these are the steps we need to take:\n\n1. Go to notebook options (on the left, below the \"Data\" explorer) and disable \"Internet\".\n2. Click \"Save Version\" (top right).\n3. Open notebook under the \"Code\" tab of the competition. It will show up under \"Your work\".\n4. Now click on the three dots in the upper right corner and select \"Submit to Competition\" (see screenshot below).\n5. Follow the on-screen instructions.\n6. Wait for the notebook to finish, results will show up under \"My Submissions\".\n\n![How to submit to the competition](https://tuc.cloud/index.php/s/z9eWEA8ZtbHki3i/preview)\n\nThat's it. Leave a comment if you have any remarks and please don't hesitate to start a new forum thread if you have any questions.","metadata":{}}]}