{
  "id": 307941,
  "title": "Meet the hosts",
  "url": "/competitions/birdclef-2022/discussion/307941",
  "author_name": "Stefan Kahl",
  "post_date": "2022-02-16T10:18:12.462000",
  "votes": 21,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thank you to everyone who is participating in this competition. Your contribution will benefit ongoing research to protect endangered bird species.</p>\n<p>As hosts of the competition, we will try to be as active and responsive as possible to assist you in your endeavor. And all without giving away any secrets about the test data, so don't bother asking :)</p>\n<p>In this thread, I will give all the hosts a chance to introduce themselves. I will be the first to do so.</p>\n<p>I am a research associate within the K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, focusing on the development of advanced machine learning models for automatic detection and identification of bird species in large audio collections. Furthermore, I am the technology lead for the <a href=\"https://birdnet.cornell.edu\" target=\"_blank\">BirdNET project</a> and have been organizing the BirdCLEF Challenge since 2018. Feel free to ask me anything related to Deep Learning for bioacoustics, I might be able to help you.</p>",
  "messages": [
    {
      "id": 1692882,
      "postDate": "2022-02-16T10:18:12.463Z",
      "content": "<p>Thank you to everyone who is participating in this competition. Your contribution will benefit ongoing research to protect endangered bird species.</p>\n<p>As hosts of the competition, we will try to be as active and responsive as possible to assist you in your endeavor. And all without giving away any secrets about the test data, so don't bother asking :)</p>\n<p>In this thread, I will give all the hosts a chance to introduce themselves. I will be the first to do so.</p>\n<p>I am a research associate within the K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, focusing on the development of advanced machine learning models for automatic detection and identification of bird species in large audio collections. Furthermore, I am the technology lead for the <a href=\"https://birdnet.cornell.edu\" target=\"_blank\">BirdNET project</a> and have been organizing the BirdCLEF Challenge since 2018. Feel free to ask me anything related to Deep Learning for bioacoustics, I might be able to help you.</p>",
      "rawMarkdown": "Thank you to everyone who is participating in this competition. Your contribution will benefit ongoing research to protect endangered bird species.\n\nAs hosts of the competition, we will try to be as active and responsive as possible to assist you in your endeavor. And all without giving away any secrets about the test data, so don't bother asking :)\n\nIn this thread, I will give all the hosts a chance to introduce themselves. I will be the first to do so.\n\nI am a research associate within the K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, focusing on the development of advanced machine learning models for automatic detection and identification of bird species in large audio collections. Furthermore, I am the technology lead for the [BirdNET project](https://birdnet.cornell.edu) and have been organizing the BirdCLEF Challenge since 2018. Feel free to ask me anything related to Deep Learning for bioacoustics, I might be able to help you.",
      "votes": 21
    },
    {
      "id": 1693618,
      "postDate": "2022-02-16T20:02:15.393Z",
      "content": "<p>Hi, all!</p>\n<p>I'm back as a co-organizer, and will be around to answer the occasional question.</p>\n<p>I've been at Google for n+2 years. I've been a 20%'er on bioacoustics for the last three-ish years, often working with the Cornell Lab on organizing these competitions, but also working with the California Academy of Sciences to understand post-fire habitat recovery after prescribed fires. I was recently a coauthor on a paper using source separation to improve birdsong classification. (Links: <a href=\"https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html\" target=\"_blank\">AIBlog</a>, <a href=\"https://arxiv.org/abs/2110.03209\" target=\"_blank\">arxiv</a>, <a href=\"https://github.com/google-research/sound-separation/tree/master/models/bird_mixit\" target=\"_blank\">github separation model</a>, <a href=\"https://bird-mixit.github.io/\" target=\"_blank\">github examples</a>.)</p>\n<p>In my day job, I work on low-bitrate audio compression (such as <a href=\"https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-codec-for.html\" target=\"_blank\">Lyra</a> and <a href=\"https://ai.googleblog.com/2021/08/soundstream-end-to-end-neural-audio.html\" target=\"_blank\">Soundstream</a>). Mostly I find new modeling approaches to make neural audio synthesis run hella fast on-device without sacrificing audio quality.</p>\n<p>We chose this year's competition topic to reflect some real problems we see in the field: Degraded classifier performance in tropical environments, and identifying threatened species which often have scant training data. I'm looking forward to seeing what ya'll come up with; it could really end up helping save threatened species in Hawaiʻi and beyond.</p>",
      "rawMarkdown": "Hi, all!\n\nI'm back as a co-organizer, and will be around to answer the occasional question.\n\nI've been at Google for n+2 years. I've been a 20%'er on bioacoustics for the last three-ish years, often working with the Cornell Lab on organizing these competitions, but also working with the California Academy of Sciences to understand post-fire habitat recovery after prescribed fires. I was recently a coauthor on a paper using source separation to improve birdsong classification. (Links: [AIBlog](https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html), [arxiv](https://arxiv.org/abs/2110.03209), [github separation model](https://github.com/google-research/sound-separation/tree/master/models/bird_mixit), [github examples](https://bird-mixit.github.io/).)\n\nIn my day job, I work on low-bitrate audio compression (such as [Lyra](https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-codec-for.html) and [Soundstream](https://ai.googleblog.com/2021/08/soundstream-end-to-end-neural-audio.html)). Mostly I find new modeling approaches to make neural audio synthesis run hella fast on-device without sacrificing audio quality.\n\nWe chose this year's competition topic to reflect some real problems we see in the field: Degraded classifier performance in tropical environments, and identifying threatened species which often have scant training data. I'm looking forward to seeing what ya'll come up with; it could really end up helping save threatened species in Hawaiʻi and beyond.",
      "votes": 13
    },
    {
      "id": 1693696,
      "postDate": "2022-02-16T21:55:42.777Z",
      "content": "<p>Aloha everyone! I am an acoustic bioinformatic specialist with the University of Hawai'i at Hilo's Listening Observatory for Hawaiian Ecosystems (LOHE). I am a bird nerd at heart and am passionate about the conservation of endangered bird species. That's why I'm excited to be part of the BirdNET project and the BirdCLEF 2022 Challenge, the tools that we are developing are instrumental in making endangered species monitoring more logistically feasible. I would be happy to answer any of your questions about Hawaiian bird species and how the LOHE lab collects and uses bioacoustic data!</p>",
      "rawMarkdown": "Aloha everyone! I am an acoustic bioinformatic specialist with the University of Hawai'i at Hilo's Listening Observatory for Hawaiian Ecosystems (LOHE). I am a bird nerd at heart and am passionate about the conservation of endangered bird species. That's why I'm excited to be part of the BirdNET project and the BirdCLEF 2022 Challenge, the tools that we are developing are instrumental in making endangered species monitoring more logistically feasible. I would be happy to answer any of your questions about Hawaiian bird species and how the LOHE lab collects and uses bioacoustic data!",
      "votes": 10
    },
    {
      "id": 1694795,
      "postDate": "2022-02-17T17:59:44.683Z",
      "content": "<p>Hi, everyone!</p>\n<p>I am the director of the K. Lisa Yang Center for Conservation Bioacoustics @ Cornell and a co-host of this competition. I am excited that BirdCLEF 2022 is now live, and I wish everyone a very successful competition! Looking forward to learning from you and your solutions, and appreciate your contribution to the conservation of Hawaiian birds!</p>\n<p>Cheers,</p>\n<p>Holger</p>",
      "rawMarkdown": "Hi, everyone!\n\nI am the director of the K. Lisa Yang Center for Conservation Bioacoustics @ Cornell and a co-host of this competition. I am excited that BirdCLEF 2022 is now live, and I wish everyone a very successful competition! Looking forward to learning from you and your solutions, and appreciate your contribution to the conservation of Hawaiian birds!\n\nCheers,\n\nHolger",
      "votes": 6
    },
    {
      "id": 1693901,
      "postDate": "2022-02-17T03:32:11.367Z",
      "content": "<p>Helloooo! It's so great to learn more about who's behind the red colored icons.</p>\n<p>If I may ask hosts, What are your favorite birds (that may or may not be in the dataset)? :) </p>",
      "rawMarkdown": "Helloooo! It's so great to learn more about who's behind the red colored icons.\n\nIf I may ask hosts, What are your favorite birds (that may or may not be in the dataset)? :) ",
      "votes": 2,
      "replies": [
        {
          "id": 1694753,
          "postDate": "2022-02-17T17:28:32.967Z",
          "content": "<p>I get that question so often and it's incredibly challenging for me to answer every time! I would have to say one of my favorite birds in Hawai'i is the 'Elepaio, which is in my profile picture. They're curious, outgoing, and oh so adorable! Sometimes they will follow you around the forest and make calls that sound like \"wooow!\" as if they're intrigued by what you're doing. But I have to throw a shout-out to the Hawai'i 'Amakihi, who are in this competition. They're being to develop immunity to avian malaria, a huge threat to island bird life, and have become a symbol of hope for Hawaiian honeycreepers. They were the focal of my master's research for that reason! </p>",
          "rawMarkdown": "I get that question so often and it's incredibly challenging for me to answer every time! I would have to say one of my favorite birds in Hawai'i is the 'Elepaio, which is in my profile picture. They're curious, outgoing, and oh so adorable! Sometimes they will follow you around the forest and make calls that sound like \"wooow!\" as if they're intrigued by what you're doing. But I have to throw a shout-out to the Hawai'i 'Amakihi, who are in this competition. They're being to develop immunity to avian malaria, a huge threat to island bird life, and have become a symbol of hope for Hawaiian honeycreepers. They were the focal of my master's research for that reason! ",
          "votes": 5
        },
        {
          "id": 1694796,
          "postDate": "2022-02-17T18:01:42.030Z",
          "content": "<p>Mine is the Eurasian hoopoe (Upupa epops). Best scientific name ever 😂</p>",
          "rawMarkdown": "Mine is the Eurasian hoopoe (Upupa epops). Best scientific name ever 😂",
          "votes": 3
        },
        {
          "id": 1710236,
          "postDate": "2022-03-02T20:34:51.407Z",
          "content": "<p>The best bird is the bird in front of me. :) </p>\n<p>That said, I really love the sound of <a href=\"https://xeno-canto.org/692187\" target=\"_blank\">Swainson's Thrush</a>, and many of the other thrushes.</p>",
          "rawMarkdown": "The best bird is the bird in front of me. :) \n\nThat said, I really love the sound of [Swainson's Thrush](https://xeno-canto.org/692187), and many of the other thrushes.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1795650,
      "postDate": "2022-05-20T02:22:54.807Z",
      "content": "<p>Could we please have more digits in the LB score?</p>",
      "rawMarkdown": "Could we please have more digits in the LB score?"
    },
    {
      "id": 1770089,
      "postDate": "2022-04-28T00:56:55.683Z",
      "content": "<p>A question:</p>\n<p>\"Thankfully, recent advances in machine learning have made it possible to automatically identify bird songs for common species with ample training data. However, it remains challenging to develop such tools for rare and endangered species, such as those in Hawai'i.\"</p>\n<p>Why can't we use the same algorithms and just manually record or take these sounds from the library? And apply the algorithms on these sounds?</p>",
      "rawMarkdown": "A question:\n\n\"Thankfully, recent advances in machine learning have made it possible to automatically identify bird songs for common species with ample training data. However, it remains challenging to develop such tools for rare and endangered species, such as those in Hawai'i.\"\n\nWhy can't we use the same algorithms and just manually record or take these sounds from the library? And apply the algorithms on these sounds?"
    },
    {
      "id": 1751223,
      "postDate": "2022-04-10T14:13:56.863Z",
      "content": "<p>can some one tell me what is the wrong with my code, I either get 0.48 or 0.51 ?</p>\n<pre><code># ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] &gt; 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n\n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n\n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n\n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n</code></pre>\n<p>The result i get when I run</p>\n<pre><code>Number of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n</code></pre>\n<p>When I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!</p>\n<p>when I train I use  only the scored classes sounds:<br>\nI train my model on augmented dataset  and my validation set is the original data set</p>\n<p>I think my problem is from the code when I submit but I don't know where.<br>\nAny advice?</p>",
      "rawMarkdown": "can some one tell me what is the wrong with my code, I either get 0.48 or 0.51 ?\n\n\n```\n# ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] > 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n        \n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n        \n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n        \n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n```\n\nThe result i get when I run\n```\nNumber of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n```\n\nWhen I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!\n\nwhen I train I use  only the scored classes sounds:\nI train my model on augmented dataset  and my validation set is the original data set\n\nI think my problem is from the code when I submit but I don't know where.\nAny advice?"
    },
    {
      "id": 1738540,
      "postDate": "2022-03-29T11:07:32.700Z",
      "content": "<p>Hi all,</p>\n<p>I'm trying to submit the code but it's showing a scoring error, is there any mistake I'm making in the wrong interpretation.</p>\n<p>======================CODE============================</p>\n<pre><code>test_audio_dir = '../input/birdclef-2022/test_soundscapes/'\nfile_list = os.listdir(test_audio_dir)\nfile_list = pd.DataFrame([x.split('.')[0] for x in file_list], columns=['file_name'])\n\nm = new_scored_data[['primary_label', 'primary_label_encoded']]\n\n# mapping label encod to birst species\ndict_map = {}\nfor x in m.sort_values(by='primary_label_encoded').drop_duplicates().values:\n    dict_map[x[1]] = x[0] \n\ndef file_duration_test(x):\n    _audio_file_path = os.path.join(\n        test_audio_dir,\n        F'{x}.ogg'\n    )\n    info = torchaudio.info(_audio_file_path)\n    return np.array([info.num_frames/info.sample_rate, info.sample_rate, info.num_channels])\n\nmeta_d = file_list['file_name'].apply(file_duration_test)\nfile_list['duration'] = [x[0] for x in meta_d.values]\nfile_list['sr'] = [x[1] for x in meta_d.values]\nfile_list['ch'] = [x[2] for x in meta_d.values]\nfile_list['n_chunks'] = file_list['duration'] / 5\n\nclass TestDataset(Dataset):\n    def __init__(self, dataset, filename, dr, sr, ch, n_chunks):\n        self.dataset = dataset\n        self.filename = filename\n        self.dr = dr\n        self.sr = sr\n        self.ch = ch\n        self.n_chunks = n_chunks\n        self.reshample = 16000\n        self.sec = 5\n\n    def __len__(self):\n        return len(self.dataset)\n\n    def __getitem__(self, idx):\n        data = self.dataset.iloc[idx]\n        n_cunks = 12\n        frame_per_chunk = self.reshample * self.sec\n        num_frames = data[self.sr] * n_cunks * self.sec\n        file = data[self.filename]\n        try:\n\n            path = os.path.join(\n                test_audio_dir,\n                F'{file}.ogg'\n            )\n\n            audio, rate = torchaudio.load(\n                path,\n                num_frames=int(num_frames), \n                frame_offset=0,\n            )\n\n            audio_reshample = TF.resample(\n                audio,\n                rate,\n                self.reshample,\n                lowpass_filter_width=16\n            )\n\n            del audio\n\n            return torch.reshape(\n                audio_reshample[:1], \n                (int(n_cunks), 1, -1)\n            ), file\n        except:\n            return np.zeros((n_cunks, 1, frame_per_chunk)), file\n\ntest_dataset = TestDataset(file_list, 'file_name', 'duration', 'sr', 'ch', 'n_chunks')\n\ndef perdict_test_data():\n    model.train(False)\n    pred_frame = {'row_id': [], 'target': []}\n    for idx in range(len(test_dataset)):\n        x, file = test_dataset[idx]\n        x_ = x.to(device)\n        y = model(x_)\n        y_ = nn.functional.softmax(y, dim = 1)\n        del x, x_, y\n        for idxy in range(len(y_)):\n            y__ = y_[idxy]\n            for idx in range(len(y__)):\n                chunk_end_time = (idx + 1) * 5\n                bird = dict_map[idx]\n                pred = y__[idx] &gt; 0.5\n                row_id = file + '_' + bird + '_' + str(chunk_end_time)\n                pred_frame['row_id'].append(row_id)\n                pred_frame['target'].append(True if pred else False)\n                del pred\n        del y_\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n        return pd.DataFrame(pred_frame, columns = ['row_id', 'target'])\n\nresult = perdict_test_data()\n\n\n# Quick sanity check\nprint(result.head()) \n\n# Convert our results to csv\nresult.to_csv(\"submission.csv\", index=False)\n</code></pre>\n<p>=================END=====================</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Hi all,\n\nI'm trying to submit the code but it's showing a scoring error, is there any mistake I'm making in the wrong interpretation.\n\n======================CODE============================\n```\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\nfile_list = os.listdir(test_audio_dir)\nfile_list = pd.DataFrame([x.split('.')[0] for x in file_list], columns=['file_name'])\n\nm = new_scored_data[['primary_label', 'primary_label_encoded']]\n\n# mapping label encod to birst species\ndict_map = {}\nfor x in m.sort_values(by='primary_label_encoded').drop_duplicates().values:\n    dict_map[x[1]] = x[0] \n\ndef file_duration_test(x):\n    _audio_file_path = os.path.join(\n        test_audio_dir,\n        F'{x}.ogg'\n    )\n    info = torchaudio.info(_audio_file_path)\n    return np.array([info.num_frames/info.sample_rate, info.sample_rate, info.num_channels])\n\nmeta_d = file_list['file_name'].apply(file_duration_test)\nfile_list['duration'] = [x[0] for x in meta_d.values]\nfile_list['sr'] = [x[1] for x in meta_d.values]\nfile_list['ch'] = [x[2] for x in meta_d.values]\nfile_list['n_chunks'] = file_list['duration'] / 5\n\nclass TestDataset(Dataset):\n    def __init__(self, dataset, filename, dr, sr, ch, n_chunks):\n        self.dataset = dataset\n        self.filename = filename\n        self.dr = dr\n        self.sr = sr\n        self.ch = ch\n        self.n_chunks = n_chunks\n        self.reshample = 16000\n        self.sec = 5\n        \n    def __len__(self):\n        return len(self.dataset)\n    \n    def __getitem__(self, idx):\n        data = self.dataset.iloc[idx]\n        n_cunks = 12\n        frame_per_chunk = self.reshample * self.sec\n        num_frames = data[self.sr] * n_cunks * self.sec\n        file = data[self.filename]\n        try:\n\n            path = os.path.join(\n                test_audio_dir,\n                F'{file}.ogg'\n            )\n\n            audio, rate = torchaudio.load(\n                path,\n                num_frames=int(num_frames), \n                frame_offset=0,\n            )\n\n            audio_reshample = TF.resample(\n                audio,\n                rate,\n                self.reshample,\n                lowpass_filter_width=16\n            )\n\n            del audio\n\n            return torch.reshape(\n                audio_reshample[:1], \n                (int(n_cunks), 1, -1)\n            ), file\n        except:\n            return np.zeros((n_cunks, 1, frame_per_chunk)), file\n\ntest_dataset = TestDataset(file_list, 'file_name', 'duration', 'sr', 'ch', 'n_chunks')\n\ndef perdict_test_data():\n    model.train(False)\n    pred_frame = {'row_id': [], 'target': []}\n    for idx in range(len(test_dataset)):\n        x, file = test_dataset[idx]\n        x_ = x.to(device)\n        y = model(x_)\n        y_ = nn.functional.softmax(y, dim = 1)\n        del x, x_, y\n        for idxy in range(len(y_)):\n            y__ = y_[idxy]\n            for idx in range(len(y__)):\n                chunk_end_time = (idx + 1) * 5\n                bird = dict_map[idx]\n                pred = y__[idx] > 0.5\n                row_id = file + '_' + bird + '_' + str(chunk_end_time)\n                pred_frame['row_id'].append(row_id)\n                pred_frame['target'].append(True if pred else False)\n                del pred\n        del y_\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n        return pd.DataFrame(pred_frame, columns = ['row_id', 'target'])\n\nresult = perdict_test_data()\n\n\n# Quick sanity check\nprint(result.head()) \n    \n# Convert our results to csv\nresult.to_csv(\"submission.csv\", index=False)\n\n```\n\n=================END=====================\n\nThank you.",
      "replies": [
        {
          "id": 1738562,
          "postDate": "2022-03-29T11:21:06.857Z",
          "content": "<p>Try the following instead of<br>\nfile_list['n_chunks'] = file_list['duration'] / 5.</p>\n<p><strong>file_list['n_chunks'] = round(file_list['duration'] / 5)</strong></p>",
          "rawMarkdown": "Try the following instead of\nfile_list['n_chunks'] = file_list['duration'] / 5.\n\n**file_list['n_chunks'] = round(file_list['duration'] / 5)**"
        }
      ]
    },
    {
      "id": 1800597,
      "postDate": "2022-05-25T04:24:46.633Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1765769,
      "postDate": "2022-04-23T20:26:32.080Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1693618,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2022-02-16T20:02:15.393000",
      "content": "<p>Hi, all!</p>\n<p>I'm back as a co-organizer, and will be around to answer the occasional question.</p>\n<p>I've been at Google for n+2 years. I've been a 20%'er on bioacoustics for the last three-ish years, often working with the Cornell Lab on organizing these competitions, but also working with the California Academy of Sciences to understand post-fire habitat recovery after prescribed fires. I was recently a coauthor on a paper using source separation to improve birdsong classification. (Links: <a href=\"https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html\" target=\"_blank\">AIBlog</a>, <a href=\"https://arxiv.org/abs/2110.03209\" target=\"_blank\">arxiv</a>, <a href=\"https://github.com/google-research/sound-separation/tree/master/models/bird_mixit\" target=\"_blank\">github separation model</a>, <a href=\"https://bird-mixit.github.io/\" target=\"_blank\">github examples</a>.)</p>\n<p>In my day job, I work on low-bitrate audio compression (such as <a href=\"https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-codec-for.html\" target=\"_blank\">Lyra</a> and <a href=\"https://ai.googleblog.com/2021/08/soundstream-end-to-end-neural-audio.html\" target=\"_blank\">Soundstream</a>). Mostly I find new modeling approaches to make neural audio synthesis run hella fast on-device without sacrificing audio quality.</p>\n<p>We chose this year's competition topic to reflect some real problems we see in the field: Degraded classifier performance in tropical environments, and identifying threatened species which often have scant training data. I'm looking forward to seeing what ya'll come up with; it could really end up helping save threatened species in Hawaiʻi and beyond.</p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 1693696,
      "author_name": "Amanda Navine",
      "author_url": "",
      "post_date": "2022-02-16T21:55:42.777000",
      "content": "<p>Aloha everyone! I am an acoustic bioinformatic specialist with the University of Hawai'i at Hilo's Listening Observatory for Hawaiian Ecosystems (LOHE). I am a bird nerd at heart and am passionate about the conservation of endangered bird species. That's why I'm excited to be part of the BirdNET project and the BirdCLEF 2022 Challenge, the tools that we are developing are instrumental in making endangered species monitoring more logistically feasible. I would be happy to answer any of your questions about Hawaiian bird species and how the LOHE lab collects and uses bioacoustic data!</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 1694795,
      "author_name": "Holger Klinck",
      "author_url": "",
      "post_date": "2022-02-17T17:59:44.683000",
      "content": "<p>Hi, everyone!</p>\n<p>I am the director of the K. Lisa Yang Center for Conservation Bioacoustics @ Cornell and a co-host of this competition. I am excited that BirdCLEF 2022 is now live, and I wish everyone a very successful competition! Looking forward to learning from you and your solutions, and appreciate your contribution to the conservation of Hawaiian birds!</p>\n<p>Cheers,</p>\n<p>Holger</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1693901,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2022-02-17T03:32:11.367000",
      "content": "<p>Helloooo! It's so great to learn more about who's behind the red colored icons.</p>\n<p>If I may ask hosts, What are your favorite birds (that may or may not be in the dataset)? :) </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1694753,
          "author_name": "Amanda Navine",
          "author_url": "",
          "post_date": "2022-02-17T17:28:32.967000",
          "content": "<p>I get that question so often and it's incredibly challenging for me to answer every time! I would have to say one of my favorite birds in Hawai'i is the 'Elepaio, which is in my profile picture. They're curious, outgoing, and oh so adorable! Sometimes they will follow you around the forest and make calls that sound like \"wooow!\" as if they're intrigued by what you're doing. But I have to throw a shout-out to the Hawai'i 'Amakihi, who are in this competition. They're being to develop immunity to avian malaria, a huge threat to island bird life, and have become a symbol of hope for Hawaiian honeycreepers. They were the focal of my master's research for that reason! </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1694796,
          "author_name": "Holger Klinck",
          "author_url": "",
          "post_date": "2022-02-17T18:01:42.030000",
          "content": "<p>Mine is the Eurasian hoopoe (Upupa epops). Best scientific name ever 😂</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1710236,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2022-03-02T20:34:51.407000",
          "content": "<p>The best bird is the bird in front of me. :) </p>\n<p>That said, I really love the sound of <a href=\"https://xeno-canto.org/692187\" target=\"_blank\">Swainson's Thrush</a>, and many of the other thrushes.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1795650,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "2022-05-20T02:22:54.807000",
      "content": "<p>Could we please have more digits in the LB score?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1770089,
      "author_name": "meenadpmd",
      "author_url": "",
      "post_date": "2022-04-28T00:56:55.683000",
      "content": "<p>A question:</p>\n<p>\"Thankfully, recent advances in machine learning have made it possible to automatically identify bird songs for common species with ample training data. However, it remains challenging to develop such tools for rare and endangered species, such as those in Hawai'i.\"</p>\n<p>Why can't we use the same algorithms and just manually record or take these sounds from the library? And apply the algorithms on these sounds?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1751223,
      "author_name": "Bahaa al-deen Kattan",
      "author_url": "",
      "post_date": "2022-04-10T14:13:56.863000",
      "content": "<p>can some one tell me what is the wrong with my code, I either get 0.48 or 0.51 ?</p>\n<pre><code># ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] &gt; 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n\n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n\n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n\n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n</code></pre>\n<p>The result i get when I run</p>\n<pre><code>Number of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n</code></pre>\n<p>When I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!</p>\n<p>when I train I use  only the scored classes sounds:<br>\nI train my model on augmented dataset  and my validation set is the original data set</p>\n<p>I think my problem is from the code when I submit but I don't know where.<br>\nAny advice?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1738540,
      "author_name": "Shubham Srivastava",
      "author_url": "",
      "post_date": "2022-03-29T11:07:32.700000",
      "content": "<p>Hi all,</p>\n<p>I'm trying to submit the code but it's showing a scoring error, is there any mistake I'm making in the wrong interpretation.</p>\n<p>======================CODE============================</p>\n<pre><code>test_audio_dir = '../input/birdclef-2022/test_soundscapes/'\nfile_list = os.listdir(test_audio_dir)\nfile_list = pd.DataFrame([x.split('.')[0] for x in file_list], columns=['file_name'])\n\nm = new_scored_data[['primary_label', 'primary_label_encoded']]\n\n# mapping label encod to birst species\ndict_map = {}\nfor x in m.sort_values(by='primary_label_encoded').drop_duplicates().values:\n    dict_map[x[1]] = x[0] \n\ndef file_duration_test(x):\n    _audio_file_path = os.path.join(\n        test_audio_dir,\n        F'{x}.ogg'\n    )\n    info = torchaudio.info(_audio_file_path)\n    return np.array([info.num_frames/info.sample_rate, info.sample_rate, info.num_channels])\n\nmeta_d = file_list['file_name'].apply(file_duration_test)\nfile_list['duration'] = [x[0] for x in meta_d.values]\nfile_list['sr'] = [x[1] for x in meta_d.values]\nfile_list['ch'] = [x[2] for x in meta_d.values]\nfile_list['n_chunks'] = file_list['duration'] / 5\n\nclass TestDataset(Dataset):\n    def __init__(self, dataset, filename, dr, sr, ch, n_chunks):\n        self.dataset = dataset\n        self.filename = filename\n        self.dr = dr\n        self.sr = sr\n        self.ch = ch\n        self.n_chunks = n_chunks\n        self.reshample = 16000\n        self.sec = 5\n\n    def __len__(self):\n        return len(self.dataset)\n\n    def __getitem__(self, idx):\n        data = self.dataset.iloc[idx]\n        n_cunks = 12\n        frame_per_chunk = self.reshample * self.sec\n        num_frames = data[self.sr] * n_cunks * self.sec\n        file = data[self.filename]\n        try:\n\n            path = os.path.join(\n                test_audio_dir,\n                F'{file}.ogg'\n            )\n\n            audio, rate = torchaudio.load(\n                path,\n                num_frames=int(num_frames), \n                frame_offset=0,\n            )\n\n            audio_reshample = TF.resample(\n                audio,\n                rate,\n                self.reshample,\n                lowpass_filter_width=16\n            )\n\n            del audio\n\n            return torch.reshape(\n                audio_reshample[:1], \n                (int(n_cunks), 1, -1)\n            ), file\n        except:\n            return np.zeros((n_cunks, 1, frame_per_chunk)), file\n\ntest_dataset = TestDataset(file_list, 'file_name', 'duration', 'sr', 'ch', 'n_chunks')\n\ndef perdict_test_data():\n    model.train(False)\n    pred_frame = {'row_id': [], 'target': []}\n    for idx in range(len(test_dataset)):\n        x, file = test_dataset[idx]\n        x_ = x.to(device)\n        y = model(x_)\n        y_ = nn.functional.softmax(y, dim = 1)\n        del x, x_, y\n        for idxy in range(len(y_)):\n            y__ = y_[idxy]\n            for idx in range(len(y__)):\n                chunk_end_time = (idx + 1) * 5\n                bird = dict_map[idx]\n                pred = y__[idx] &gt; 0.5\n                row_id = file + '_' + bird + '_' + str(chunk_end_time)\n                pred_frame['row_id'].append(row_id)\n                pred_frame['target'].append(True if pred else False)\n                del pred\n        del y_\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n        return pd.DataFrame(pred_frame, columns = ['row_id', 'target'])\n\nresult = perdict_test_data()\n\n\n# Quick sanity check\nprint(result.head()) \n\n# Convert our results to csv\nresult.to_csv(\"submission.csv\", index=False)\n</code></pre>\n<p>=================END=====================</p>\n<p>Thank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1738562,
          "author_name": "Attila Ambrus",
          "author_url": "",
          "post_date": "2022-03-29T11:21:06.857000",
          "content": "<p>Try the following instead of<br>\nfile_list['n_chunks'] = file_list['duration'] / 5.</p>\n<p><strong>file_list['n_chunks'] = round(file_list['duration'] / 5)</strong></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1800597,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-25T04:24:46.633000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1765769,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T20:26:32.080000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1692882": "Thank you to everyone who is participating in this competition. Your contribution will benefit ongoing research to protect endangered bird species.\n\nAs hosts of the competition, we will try to be as active and responsive as possible to assist you in your endeavor. And all without giving away any secrets about the test data, so don't bother asking :)\n\nIn this thread, I will give all the hosts a chance to introduce themselves. I will be the first to do so.\n\nI am a research associate within the K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology, focusing on the development of advanced machine learning models for automatic detection and identification of bird species in large audio collections. Furthermore, I am the technology lead for the [BirdNET project](https://birdnet.cornell.edu) and have been organizing the BirdCLEF Challenge since 2018. Feel free to ask me anything related to Deep Learning for bioacoustics, I might be able to help you.",
    "1693618": "Hi, all!\n\nI'm back as a co-organizer, and will be around to answer the occasional question.\n\nI've been at Google for n+2 years. I've been a 20%'er on bioacoustics for the last three-ish years, often working with the Cornell Lab on organizing these competitions, but also working with the California Academy of Sciences to understand post-fire habitat recovery after prescribed fires. I was recently a coauthor on a paper using source separation to improve birdsong classification. (Links: [AIBlog](https://ai.googleblog.com/2022/01/separating-birdsong-in-wild-for.html), [arxiv](https://arxiv.org/abs/2110.03209), [github separation model](https://github.com/google-research/sound-separation/tree/master/models/bird_mixit), [github examples](https://bird-mixit.github.io/).)\n\nIn my day job, I work on low-bitrate audio compression (such as [Lyra](https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-codec-for.html) and [Soundstream](https://ai.googleblog.com/2021/08/soundstream-end-to-end-neural-audio.html)). Mostly I find new modeling approaches to make neural audio synthesis run hella fast on-device without sacrificing audio quality.\n\nWe chose this year's competition topic to reflect some real problems we see in the field: Degraded classifier performance in tropical environments, and identifying threatened species which often have scant training data. I'm looking forward to seeing what ya'll come up with; it could really end up helping save threatened species in Hawaiʻi and beyond.",
    "1693696": "Aloha everyone! I am an acoustic bioinformatic specialist with the University of Hawai'i at Hilo's Listening Observatory for Hawaiian Ecosystems (LOHE). I am a bird nerd at heart and am passionate about the conservation of endangered bird species. That's why I'm excited to be part of the BirdNET project and the BirdCLEF 2022 Challenge, the tools that we are developing are instrumental in making endangered species monitoring more logistically feasible. I would be happy to answer any of your questions about Hawaiian bird species and how the LOHE lab collects and uses bioacoustic data!",
    "1694795": "Hi, everyone!\n\nI am the director of the K. Lisa Yang Center for Conservation Bioacoustics @ Cornell and a co-host of this competition. I am excited that BirdCLEF 2022 is now live, and I wish everyone a very successful competition! Looking forward to learning from you and your solutions, and appreciate your contribution to the conservation of Hawaiian birds!\n\nCheers,\n\nHolger",
    "1693901": "Helloooo! It's so great to learn more about who's behind the red colored icons.\n\nIf I may ask hosts, What are your favorite birds (that may or may not be in the dataset)? :) ",
    "1795650": "Could we please have more digits in the LB score?",
    "1770089": "A question:\n\n\"Thankfully, recent advances in machine learning have made it possible to automatically identify bird songs for common species with ample training data. However, it remains challenging to develop such tools for rare and endangered species, such as those in Hawai'i.\"\n\nWhy can't we use the same algorithms and just manually record or take these sounds from the library? And apply the algorithms on these sounds?",
    "1751223": "can some one tell me what is the wrong with my code, I either get 0.48 or 0.51 ?\n\n\n```\n# ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] > 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n        \n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n        \n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n        \n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n```\n\nThe result i get when I run\n```\nNumber of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n```\n\nWhen I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!\n\nwhen I train I use  only the scored classes sounds:\nI train my model on augmented dataset  and my validation set is the original data set\n\nI think my problem is from the code when I submit but I don't know where.\nAny advice?",
    "1738540": "Hi all,\n\nI'm trying to submit the code but it's showing a scoring error, is there any mistake I'm making in the wrong interpretation.\n\n======================CODE============================\n```\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\nfile_list = os.listdir(test_audio_dir)\nfile_list = pd.DataFrame([x.split('.')[0] for x in file_list], columns=['file_name'])\n\nm = new_scored_data[['primary_label', 'primary_label_encoded']]\n\n# mapping label encod to birst species\ndict_map = {}\nfor x in m.sort_values(by='primary_label_encoded').drop_duplicates().values:\n    dict_map[x[1]] = x[0] \n\ndef file_duration_test(x):\n    _audio_file_path = os.path.join(\n        test_audio_dir,\n        F'{x}.ogg'\n    )\n    info = torchaudio.info(_audio_file_path)\n    return np.array([info.num_frames/info.sample_rate, info.sample_rate, info.num_channels])\n\nmeta_d = file_list['file_name'].apply(file_duration_test)\nfile_list['duration'] = [x[0] for x in meta_d.values]\nfile_list['sr'] = [x[1] for x in meta_d.values]\nfile_list['ch'] = [x[2] for x in meta_d.values]\nfile_list['n_chunks'] = file_list['duration'] / 5\n\nclass TestDataset(Dataset):\n    def __init__(self, dataset, filename, dr, sr, ch, n_chunks):\n        self.dataset = dataset\n        self.filename = filename\n        self.dr = dr\n        self.sr = sr\n        self.ch = ch\n        self.n_chunks = n_chunks\n        self.reshample = 16000\n        self.sec = 5\n        \n    def __len__(self):\n        return len(self.dataset)\n    \n    def __getitem__(self, idx):\n        data = self.dataset.iloc[idx]\n        n_cunks = 12\n        frame_per_chunk = self.reshample * self.sec\n        num_frames = data[self.sr] * n_cunks * self.sec\n        file = data[self.filename]\n        try:\n\n            path = os.path.join(\n                test_audio_dir,\n                F'{file}.ogg'\n            )\n\n            audio, rate = torchaudio.load(\n                path,\n                num_frames=int(num_frames), \n                frame_offset=0,\n            )\n\n            audio_reshample = TF.resample(\n                audio,\n                rate,\n                self.reshample,\n                lowpass_filter_width=16\n            )\n\n            del audio\n\n            return torch.reshape(\n                audio_reshample[:1], \n                (int(n_cunks), 1, -1)\n            ), file\n        except:\n            return np.zeros((n_cunks, 1, frame_per_chunk)), file\n\ntest_dataset = TestDataset(file_list, 'file_name', 'duration', 'sr', 'ch', 'n_chunks')\n\ndef perdict_test_data():\n    model.train(False)\n    pred_frame = {'row_id': [], 'target': []}\n    for idx in range(len(test_dataset)):\n        x, file = test_dataset[idx]\n        x_ = x.to(device)\n        y = model(x_)\n        y_ = nn.functional.softmax(y, dim = 1)\n        del x, x_, y\n        for idxy in range(len(y_)):\n            y__ = y_[idxy]\n            for idx in range(len(y__)):\n                chunk_end_time = (idx + 1) * 5\n                bird = dict_map[idx]\n                pred = y__[idx] > 0.5\n                row_id = file + '_' + bird + '_' + str(chunk_end_time)\n                pred_frame['row_id'].append(row_id)\n                pred_frame['target'].append(True if pred else False)\n                del pred\n        del y_\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n        return pd.DataFrame(pred_frame, columns = ['row_id', 'target'])\n\nresult = perdict_test_data()\n\n\n# Quick sanity check\nprint(result.head()) \n    \n# Convert our results to csv\nresult.to_csv(\"submission.csv\", index=False)\n\n```\n\n=================END=====================\n\nThank you.",
    "1800597": "",
    "1765769": ""
  }
}