{
  "id": 323344,
  "title": "How can I submit the prediction?",
  "url": "/competitions/birdclef-2022/discussion/323344",
  "author_name": "",
  "post_date": "2022-05-06T05:44:42.393007100Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello, I have done with the training but when I want to make the prediction with the testing (test_soundscapes), only one audio file from the test_soundscapes is available. Therefore I cannot make the prediction from all the file in test_soundscapes and I cannot output the submission file for submission. I used the code below to do the prediction. Any help please </p>\n<p>test_path = \"/kaggle/input/birdclef-2022/test_soundscapes/\"<br>\nfiles = [f.split('.')[0] for f in sorted(os.listdir(test_path))]</p>\n<p>birds_path = \"/kaggle/input/birdclef-2022/scored_birds.json\"<br>\nwith open(birds_path) as bf:<br>\n    birds = json.load(bf)</p>\n<p>data = []<br>\nfor f in files:<br>\n    file_path = test_path + f + '.ogg'<br>\n    audio, sr = librosa.load(file_path)<br>\n    # Get number of samples for 5 seconds; replace 5 by any number<br>\n    buffer = 5 * sr<br>\n    samples_total = len(audio)<br>\n    samples_wrote = 0<br>\n    counter = 1</p>\n<pre><code>while samples_wrote &lt; samples_total:\n    #check if the buffer is not exceeding total samples \n    if buffer &gt; (samples_total - samples_wrote):\n        buffer = samples_total - samples_wrote\n\n    block = audio[samples_wrote : (samples_wrote + buffer)]\n    feat = extractFeatures(block, sr)\n    pred = model.predict(feat)\n    label_index = np.argmax(pred,axis=1)[0]\n\n    for b in birds:\n        segment_end = counter * 5   \n        row_id = f + '_' + b + '_' + str(segment_end)\n        target = False\n        if labels[label_index] == b:\n            target = True\n        data.append([row_id, target])\n    counter += 1\n    samples_wrote += buffer\n</code></pre>\n<p>submission_df = pd.DataFrame(data, columns=['row_id', 'target'])<br>\nsubmission_df</p>",
  "messages": [
    {
      "id": "1779105",
      "postDate": "05/06/2022 05:44:42",
      "content": "<p>Hello, I have done with the training but when I want to make the prediction with the testing (test_soundscapes), only one audio file from the test_soundscapes is available. Therefore I cannot make the prediction from all the file in test_soundscapes and I cannot output the submission file for submission. I used the code below to do the prediction. Any help please </p>\n<p>test_path = \"/kaggle/input/birdclef-2022/test_soundscapes/\"<br>\nfiles = [f.split('.')[0] for f in sorted(os.listdir(test_path))]</p>\n<p>birds_path = \"/kaggle/input/birdclef-2022/scored_birds.json\"<br>\nwith open(birds_path) as bf:<br>\n    birds = json.load(bf)</p>\n<p>data = []<br>\nfor f in files:<br>\n    file_path = test_path + f + '.ogg'<br>\n    audio, sr = librosa.load(file_path)<br>\n    # Get number of samples for 5 seconds; replace 5 by any number<br>\n    buffer = 5 * sr<br>\n    samples_total = len(audio)<br>\n    samples_wrote = 0<br>\n    counter = 1</p>\n<pre><code>while samples_wrote &lt; samples_total:\n    #check if the buffer is not exceeding total samples \n    if buffer &gt; (samples_total - samples_wrote):\n        buffer = samples_total - samples_wrote\n\n    block = audio[samples_wrote : (samples_wrote + buffer)]\n    feat = extractFeatures(block, sr)\n    pred = model.predict(feat)\n    label_index = np.argmax(pred,axis=1)[0]\n\n    for b in birds:\n        segment_end = counter * 5   \n        row_id = f + '_' + b + '_' + str(segment_end)\n        target = False\n        if labels[label_index] == b:\n            target = True\n        data.append([row_id, target])\n    counter += 1\n    samples_wrote += buffer\n</code></pre>\n<p>submission_df = pd.DataFrame(data, columns=['row_id', 'target'])<br>\nsubmission_df</p>",
      "rawMarkdown": "Hello, I have done with the training but when I want to make the prediction with the testing (test_soundscapes), only one audio file from the test_soundscapes is available. Therefore I cannot make the prediction from all the file in test_soundscapes and I cannot output the submission file for submission. I used the code below to do the prediction. Any help please \n\ntest_path = \"/kaggle/input/birdclef-2022/test_soundscapes/\"\nfiles = [f.split('.')[0] for f in sorted(os.listdir(test_path))]\n\nbirds_path = \"/kaggle/input/birdclef-2022/scored_birds.json\"\nwith open(birds_path) as bf:\n    birds = json.load(bf)\n\ndata = []\nfor f in files:\n    file_path = test_path + f + '.ogg'\n    audio, sr = librosa.load(file_path)\n    # Get number of samples for 5 seconds; replace 5 by any number\n    buffer = 5 * sr\n    samples_total = len(audio)\n    samples_wrote = 0\n    counter = 1\n\n    while samples_wrote < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_wrote):\n            buffer = samples_total - samples_wrote\n\n        block = audio[samples_wrote : (samples_wrote + buffer)]\n        feat = extractFeatures(block, sr)\n        pred = model.predict(feat)\n        label_index = np.argmax(pred,axis=1)[0]\n        \n        for b in birds:\n            segment_end = counter * 5   \n            row_id = f + '_' + b + '_' + str(segment_end)\n            target = False\n            if labels[label_index] == b:\n                target = True\n            data.append([row_id, target])\n        counter += 1\n        samples_wrote += buffer\n        \nsubmission_df = pd.DataFrame(data, columns=['row_id', 'target'])\nsubmission_df",
      "votes": null
    },
    {
      "id": "1779448",
      "postDate": "05/06/2022 13:02:13",
      "content": "<p>The host prepared a manual \"How to submit to BirdCLEF 2022\". The link to the code is:<br>\n<a href=\"https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022\" target=\"_blank\">https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022</a><br>\nHope it helps.</p>",
      "rawMarkdown": "The host prepared a manual \"How to submit to BirdCLEF 2022\". The link to the code is:\nhttps://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022\nHope it helps.",
      "votes": null
    },
    {
      "id": "1779904",
      "postDate": "05/06/2022 23:41:13",
      "content": "<p>Thank you very much  </p>",
      "rawMarkdown": "Thank you very much",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1779448,
      "author_name": "blankaf",
      "author_url": "",
      "post_date": "05/06/2022 13:02:13",
      "content": "<p>The host prepared a manual \"How to submit to BirdCLEF 2022\". The link to the code is:<br>\n<a href=\"https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022\" target=\"_blank\">https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022</a><br>\nHope it helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1779904,
          "author_name": "camille1234",
          "author_url": "",
          "post_date": "05/06/2022 23:41:13",
          "content": "<p>Thank you very much  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1779105": "Hello, I have done with the training but when I want to make the prediction with the testing (test_soundscapes), only one audio file from the test_soundscapes is available. Therefore I cannot make the prediction from all the file in test_soundscapes and I cannot output the submission file for submission. I used the code below to do the prediction. Any help please \n\ntest_path = \"/kaggle/input/birdclef-2022/test_soundscapes/\"\nfiles = [f.split('.')[0] for f in sorted(os.listdir(test_path))]\n\nbirds_path = \"/kaggle/input/birdclef-2022/scored_birds.json\"\nwith open(birds_path) as bf:\n    birds = json.load(bf)\n\ndata = []\nfor f in files:\n    file_path = test_path + f + '.ogg'\n    audio, sr = librosa.load(file_path)\n    # Get number of samples for 5 seconds; replace 5 by any number\n    buffer = 5 * sr\n    samples_total = len(audio)\n    samples_wrote = 0\n    counter = 1\n\n    while samples_wrote < samples_total:\n        #check if the buffer is not exceeding total samples \n        if buffer > (samples_total - samples_wrote):\n            buffer = samples_total - samples_wrote\n\n        block = audio[samples_wrote : (samples_wrote + buffer)]\n        feat = extractFeatures(block, sr)\n        pred = model.predict(feat)\n        label_index = np.argmax(pred,axis=1)[0]\n        \n        for b in birds:\n            segment_end = counter * 5   \n            row_id = f + '_' + b + '_' + str(segment_end)\n            target = False\n            if labels[label_index] == b:\n                target = True\n            data.append([row_id, target])\n        counter += 1\n        samples_wrote += buffer\n        \nsubmission_df = pd.DataFrame(data, columns=['row_id', 'target'])\nsubmission_df",
    "1779448": "The host prepared a manual \"How to submit to BirdCLEF 2022\". The link to the code is:\nhttps://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022\nHope it helps.",
    "1779904": "Thank you very much"
  },
  "source": "meta"
}