{
  "id": 317924,
  "title": "can some one tell me what is the wrong with my code, I either get 0.48 or 0.51 ?",
  "url": "/competitions/birdclef-2022/discussion/317924",
  "author_name": "Bahaa al-deen Kattan",
  "post_date": "2022-04-09T16:08:09.280000",
  "votes": 2,
  "comment_count": 16,
  "views": 0,
  "content": "<pre><code># ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] &gt; 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n\n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n\n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n\n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n</code></pre>\n<p>The result i get when I run</p>\n<pre><code>Number of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n</code></pre>\n<p>When I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!</p>\n<p>when I train I use  only the scored classes sounds:<br>\nI train my model on augmented dataset  and my validation set is the original data set</p>\n<p>I think my problem is from the code when I submit but I don't know where.<br>\nAny advice?</p>",
  "messages": [
    {
      "id": 1750412,
      "postDate": "2022-04-09T16:08:09.280Z",
      "content": "<pre><code># ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] &gt; 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n\n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n\n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n\n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n</code></pre>\n<p>The result i get when I run</p>\n<pre><code>Number of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n</code></pre>\n<p>When I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!</p>\n<p>when I train I use  only the scored classes sounds:<br>\nI train my model on augmented dataset  and my validation set is the original data set</p>\n<p>I think my problem is from the code when I submit but I don't know where.<br>\nAny advice?</p>",
      "rawMarkdown": "```\n# ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] > 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n        \n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n        \n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n        \n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n```\n\nThe result i get when I run\n```\nNumber of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n```\n\nWhen I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!\n\nwhen I train I use  only the scored classes sounds:\nI train my model on augmented dataset  and my validation set is the original data set\n\nI think my problem is from the code when I submit but I don't know where.\nAny advice?",
      "votes": 2
    },
    {
      "id": 1751110,
      "postDate": "2022-04-10T11:31:47.517Z",
      "content": "<p>I'm actually struggling to get a correct submission too… It's been driving me mad for three days now ^^</p>\n<p>Why do you set the sample_rate to 22kHz ? Shouldn't it be 32kHz ?</p>",
      "rawMarkdown": "I'm actually struggling to get a correct submission too... It's been driving me mad for three days now ^^\n\nWhy do you set the sample_rate to 22kHz ? Shouldn't it be 32kHz ?",
      "votes": 1,
      "replies": [
        {
          "id": 1751139,
          "postDate": "2022-04-10T12:16:41.490Z",
          "content": "<p>Sample rate is the number of samples of audio carried per second,<br>\nno it's normal, you can put any variable, but notice that it will effect size (Mb) of data. <br>\nthis number does not effect on sound quality alot, for example,  most sounds are in 44kHz<br>\nbut in machine learning you can reduce this number, I think the lowest sample rate you can reduce is 16kHz</p>",
          "rawMarkdown": "Sample rate is the number of samples of audio carried per second,\nno it's normal, you can put any variable, but notice that it will effect size (Mb) of data. \nthis number does not effect on sound quality alot, for example,  most sounds are in 44kHz\nbut in machine learning you can reduce this number, I think the lowest sample rate you can reduce is 16kHz"
        },
        {
          "id": 1751140,
          "postDate": "2022-04-10T12:18:17.040Z",
          "content": "<p>Sure I was just thinking that it might be a typo, screwing with your model.</p>",
          "rawMarkdown": "Sure I was just thinking that it might be a typo, screwing with your model."
        },
        {
          "id": 1751368,
          "postDate": "2022-04-10T17:11:15.847Z",
          "content": "<p>It may be worth double checking that your resampling works well for a few different sample rates. I would try resampling and playing back audio from a couple-few source sample rates, and check that the extracted mel-spectrogram looks reasonable regardless of the source sample rate.</p>\n<p>It turns out that 22kHz is often /actually/ 22050 Hz, half of 44.1 kHz, which is the usual CD sampling rate. One thing to be careful of with sample rate: The highest frequency detectable at a given sample rate is half the sample rate (the nyquist theorem). There are birds which have meaningful non-redundant information north of 10 kHz range, so 16kHz sample rate can miss some interesting info. (Golden crowned kinglet and brown creeper are the two which come to mind for California, at least.)</p>",
          "rawMarkdown": "It may be worth double checking that your resampling works well for a few different sample rates. I would try resampling and playing back audio from a couple-few source sample rates, and check that the extracted mel-spectrogram looks reasonable regardless of the source sample rate.\n\nIt turns out that 22kHz is often /actually/ 22050 Hz, half of 44.1 kHz, which is the usual CD sampling rate. One thing to be careful of with sample rate: The highest frequency detectable at a given sample rate is half the sample rate (the nyquist theorem). There are birds which have meaningful non-redundant information north of 10 kHz range, so 16kHz sample rate can miss some interesting info. (Golden crowned kinglet and brown creeper are the two which come to mind for California, at least.)",
          "votes": 5
        }
      ]
    },
    {
      "id": 1750616,
      "postDate": "2022-04-09T21:20:20.363Z",
      "content": "<p>0.48 - all False<br>\n0.51 - all True</p>",
      "rawMarkdown": "0.48 - all False\n0.51 - all True",
      "votes": 1
    },
    {
      "id": 1751179,
      "postDate": "2022-04-10T13:08:56.293Z",
      "content": "<p>Me too. I trained a good model .f1 =0.88. But only 0.49 on leaderboard?? I have predicted all pigeons (scoring and unscoring). Gaussian noise is also added .<br>\nBut it's bad…</p>",
      "rawMarkdown": "Me too. I trained a good model .f1 =0.88. But only 0.49 on leaderboard?? I have predicted all pigeons (scoring and unscoring). Gaussian noise is also added .\nBut it's bad...",
      "replies": [
        {
          "id": 1751324,
          "postDate": "2022-04-10T15:54:49.303Z",
          "content": "<p>you seem to have a decent score to me!</p>\n<p>Did you just solve something?</p>",
          "rawMarkdown": "you seem to have a decent score to me!\n\nDid you just solve something?"
        },
        {
          "id": 1751535,
          "postDate": "2022-04-10T22:49:00.087Z",
          "content": "<p>This is why I got 0.71 points: <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318081\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/318081</a></p>\n<p>But this is not what I want.I might get some reason in it to change my code.<br>\nI got 0.51 score in my model just. <br>\nFirst I used data augmentation . The two samples are added together on the time spectrum.<br>\n(only fast Fourier transform is done).<br>\nsecondly trained many epoch (I forgot how many)<br>\nThen I removed the unscored bird from the original trained model and train one epoch for the scored birds. get F1=0.77 (only scored birds ).<br>\nI think it should be overfitting. It may be because my training set and test set are not completely divided.<br>\nI divided an audio into many 20 seconds.Then divide the training set and test set from it.<br>\nProbably the model learned something weird.But if you don't, some training samples will be few.</p>",
          "rawMarkdown": "This is why I got 0.71 points: https://www.kaggle.com/competitions/birdclef-2022/discussion/318081\n\nBut this is not what I want.I might get some reason in it to change my code.\nI got 0.51 score in my model just. \nFirst I used data augmentation . The two samples are added together on the time spectrum.\n(only fast Fourier transform is done).\nsecondly trained many epoch (I forgot how many)\nThen I removed the unscored bird from the original trained model and train one epoch for the scored birds. get F1=0.77 (only scored birds ).\nI think it should be overfitting. It may be because my training set and test set are not completely divided.\nI divided an audio into many 20 seconds.Then divide the training set and test set from it.\nProbably the model learned something weird.But if you don't, some training samples will be few."
        }
      ]
    },
    {
      "id": 1750622,
      "postDate": "2022-04-09T21:37:24.087Z",
      "content": "<p>I know but I don't know why I get this while my model can predict different sounds.<br>\nhere is my best score on my machine:</p>\n<pre><code>BEST EPOCH(5) time=(0.03)hours: loss=(train_loss=0.7492 | val_loss=0.677) | F1-score=(train_f1_score=0.568 | val_f1_score=0.61) | Accuracy=(train_accuracy=[0.77] | val_accuracy=0.793)\n</code></pre>",
      "rawMarkdown": "I know but I don't know why I get this while my model can predict different sounds.\nhere is my best score on my machine:\n```\nBEST EPOCH(5) time=(0.03)hours: loss=(train_loss=0.7492 | val_loss=0.677) | F1-score=(train_f1_score=0.568 | val_f1_score=0.61) | Accuracy=(train_accuracy=[0.77] | val_accuracy=0.793)\n\n```",
      "replies": [
        {
          "id": 1750963,
          "postDate": "2022-04-10T09:04:14.817Z",
          "content": "<p>Shared code gives you 0.48? What change gives you 0.51?</p>\n<p>Maybe order of submission file should be unchanged. You can check this notebook: <a href=\"https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline\" target=\"_blank\">https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline</a></p>",
          "rawMarkdown": "Shared code gives you 0.48? What change gives you 0.51?\n\nMaybe order of submission file should be unchanged. You can check this notebook: https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline"
        },
        {
          "id": 1751144,
          "postDate": "2022-04-10T12:22:33.953Z",
          "content": "<p>what do you mean by order, do you mean I should  predict on all birds (the scored and non-scored birds)?</p>",
          "rawMarkdown": "what do you mean by order, do you mean I should  predict on all birds (the scored and non-scored birds)?"
        },
        {
          "id": 1751148,
          "postDate": "2022-04-10T12:29:46.037Z",
          "content": "<p><a href=\"https://www.kaggle.com/hypocrites\" target=\"_blank\">@hypocrites</a> I rearranged the final submission according to <code>sample_submission.csv</code> but it does not change anything.</p>",
          "rawMarkdown": "@hypocrites I rearranged the final submission according to `sample_submission.csv` but it does not change anything."
        },
        {
          "id": 1751156,
          "postDate": "2022-04-10T12:43:43.447Z",
          "content": "<p><a href=\"https://www.kaggle.com/hypocrites\" target=\"_blank\">@hypocrites</a> can you give us an example about how you submitted to the comp??</p>",
          "rawMarkdown": "@hypocrites can you give us an example about how you submitted to the comp??"
        },
        {
          "id": 1751364,
          "postDate": "2022-04-10T17:05:49.337Z",
          "content": "<p>It does sound potentially like an issue with your submission format. Submitting correctly is a bit difficult, so we strongly recommend adapting the <a href=\"https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022/notebook\" target=\"_blank\">example submission</a>, which is 'known-good' instead of writing custom code.</p>\n<p>Looks like your code is pretty close, but worth going over it and comparing line-by-line to look for differences, or writing a couple tests for equivalence.</p>",
          "rawMarkdown": "It does sound potentially like an issue with your submission format. Submitting correctly is a bit difficult, so we strongly recommend adapting the [example submission](https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022/notebook), which is 'known-good' instead of writing custom code.\n\nLooks like your code is pretty close, but worth going over it and comparing line-by-line to look for differences, or writing a couple tests for equivalence.",
          "votes": 1
        },
        {
          "id": 1751393,
          "postDate": "2022-04-10T17:56:31.107Z",
          "content": "<p>I do it like in shared link: <br>\n1) iterate over all files in test folder and generate the dict with prediction;<br>\n2) iterate over submission file and fill it according to values inside obtained dict.</p>\n<p>Also looking on your code what I can propose to try: instead of taking just single argmax from prediction you may try take 5-10 max values.</p>",
          "rawMarkdown": "I do it like in shared link: \n1) iterate over all files in test folder and generate the dict with prediction;\n2) iterate over submission file and fill it according to values inside obtained dict.\n\nAlso looking on your code what I can propose to try: instead of taking just single argmax from prediction you may try take 5-10 max values."
        }
      ]
    },
    {
      "id": 1751132,
      "postDate": "2022-04-10T12:07:34.213Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1751110,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2022-04-10T11:31:47.517000",
      "content": "<p>I'm actually struggling to get a correct submission too… It's been driving me mad for three days now ^^</p>\n<p>Why do you set the sample_rate to 22kHz ? Shouldn't it be 32kHz ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1751139,
          "author_name": "Bahaa al-deen Kattan",
          "author_url": "",
          "post_date": "2022-04-10T12:16:41.490000",
          "content": "<p>Sample rate is the number of samples of audio carried per second,<br>\nno it's normal, you can put any variable, but notice that it will effect size (Mb) of data. <br>\nthis number does not effect on sound quality alot, for example,  most sounds are in 44kHz<br>\nbut in machine learning you can reduce this number, I think the lowest sample rate you can reduce is 16kHz</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1751140,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-04-10T12:18:17.040000",
          "content": "<p>Sure I was just thinking that it might be a typo, screwing with your model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1751368,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2022-04-10T17:11:15.847000",
          "content": "<p>It may be worth double checking that your resampling works well for a few different sample rates. I would try resampling and playing back audio from a couple-few source sample rates, and check that the extracted mel-spectrogram looks reasonable regardless of the source sample rate.</p>\n<p>It turns out that 22kHz is often /actually/ 22050 Hz, half of 44.1 kHz, which is the usual CD sampling rate. One thing to be careful of with sample rate: The highest frequency detectable at a given sample rate is half the sample rate (the nyquist theorem). There are birds which have meaningful non-redundant information north of 10 kHz range, so 16kHz sample rate can miss some interesting info. (Golden crowned kinglet and brown creeper are the two which come to mind for California, at least.)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1750616,
      "author_name": "anthony",
      "author_url": "",
      "post_date": "2022-04-09T21:20:20.363000",
      "content": "<p>0.48 - all False<br>\n0.51 - all True</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1751179,
      "author_name": "yoyobar",
      "author_url": "",
      "post_date": "2022-04-10T13:08:56.293000",
      "content": "<p>Me too. I trained a good model .f1 =0.88. But only 0.49 on leaderboard?? I have predicted all pigeons (scoring and unscoring). Gaussian noise is also added .<br>\nBut it's bad…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1751324,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-04-10T15:54:49.303000",
          "content": "<p>you seem to have a decent score to me!</p>\n<p>Did you just solve something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1751535,
          "author_name": "yoyobar",
          "author_url": "",
          "post_date": "2022-04-10T22:49:00.087000",
          "content": "<p>This is why I got 0.71 points: <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318081\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/318081</a></p>\n<p>But this is not what I want.I might get some reason in it to change my code.<br>\nI got 0.51 score in my model just. <br>\nFirst I used data augmentation . The two samples are added together on the time spectrum.<br>\n(only fast Fourier transform is done).<br>\nsecondly trained many epoch (I forgot how many)<br>\nThen I removed the unscored bird from the original trained model and train one epoch for the scored birds. get F1=0.77 (only scored birds ).<br>\nI think it should be overfitting. It may be because my training set and test set are not completely divided.<br>\nI divided an audio into many 20 seconds.Then divide the training set and test set from it.<br>\nProbably the model learned something weird.But if you don't, some training samples will be few.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1750622,
      "author_name": "Bahaa al-deen Kattan",
      "author_url": "",
      "post_date": "2022-04-09T21:37:24.087000",
      "content": "<p>I know but I don't know why I get this while my model can predict different sounds.<br>\nhere is my best score on my machine:</p>\n<pre><code>BEST EPOCH(5) time=(0.03)hours: loss=(train_loss=0.7492 | val_loss=0.677) | F1-score=(train_f1_score=0.568 | val_f1_score=0.61) | Accuracy=(train_accuracy=[0.77] | val_accuracy=0.793)\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 1750963,
          "author_name": "anthony",
          "author_url": "",
          "post_date": "2022-04-10T09:04:14.817000",
          "content": "<p>Shared code gives you 0.48? What change gives you 0.51?</p>\n<p>Maybe order of submission file should be unchanged. You can check this notebook: <a href=\"https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline\" target=\"_blank\">https://www.kaggle.com/code/tattaka/birdclef2022-submission-baseline</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1751144,
          "author_name": "Bahaa al-deen Kattan",
          "author_url": "",
          "post_date": "2022-04-10T12:22:33.953000",
          "content": "<p>what do you mean by order, do you mean I should  predict on all birds (the scored and non-scored birds)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1751148,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-04-10T12:29:46.037000",
          "content": "<p><a href=\"https://www.kaggle.com/hypocrites\" target=\"_blank\">@hypocrites</a> I rearranged the final submission according to <code>sample_submission.csv</code> but it does not change anything.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1751156,
          "author_name": "Bahaa al-deen Kattan",
          "author_url": "",
          "post_date": "2022-04-10T12:43:43.447000",
          "content": "<p><a href=\"https://www.kaggle.com/hypocrites\" target=\"_blank\">@hypocrites</a> can you give us an example about how you submitted to the comp??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1751364,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2022-04-10T17:05:49.337000",
          "content": "<p>It does sound potentially like an issue with your submission format. Submitting correctly is a bit difficult, so we strongly recommend adapting the <a href=\"https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2022/notebook\" target=\"_blank\">example submission</a>, which is 'known-good' instead of writing custom code.</p>\n<p>Looks like your code is pretty close, but worth going over it and comparing line-by-line to look for differences, or writing a couple tests for equivalence.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1751393,
          "author_name": "anthony",
          "author_url": "",
          "post_date": "2022-04-10T17:56:31.107000",
          "content": "<p>I do it like in shared link: <br>\n1) iterate over all files in test folder and generate the dict with prediction;<br>\n2) iterate over submission file and fill it according to values inside obtained dict.</p>\n<p>Also looking on your code what I can propose to try: instead of taking just single argmax from prediction you may try take 5-10 max values.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1751132,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-10T12:07:34.213000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1750412": "```\n# ---import libraries--- #\nimport os,torchvision\nimport warnings\nimport cv2\nimport librosa\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport torch, torchaudio\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_columns',20)\npd.set_option('display.width', 2000)\n\ntest_df = pd.read_csv('../input/birdclef-2022/test.csv')\nsub = pd.read_csv('../input/birdclef-2022/sample_submission.csv')\n\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'\nbest_state = torch.load('../input/best-model-noise/best_model_0.61_fold_1.pt', map_location=torch.device(device))\n\nbest_model = CNNNetwork(22)\nbest_model.load_state_dict(best_state)\nbest_model.to(device)\n\nbest_model.eval()\n\n# --- preprare data --- #\ndef _resample_if_necessary(signal, sr):\n    if sr != target_sample_rate:\n        signal = librosa.resample(signal,sr, target_sample_rate)\n    return signal\n\ndef _mix_down_if_necessary(signal):\n    if len(signal.shape) == 2:\n        if signal.shape[1] > 1:\n            signal = signal[:, 0]  # torch.mean(signal, dim=1, keepdim=True)\n    return signal\n\nnum_samples = 5\ntarget_sample_rate = 22_000\ntotal_length = target_sample_rate*num_samples\n\nnormalize = torchvision.transforms.Normalize([-74.30503278616727], [23.726224427065763])\n\ntest_audio_dir = '../input/birdclef-2022/test_soundscapes/'\n\ndict_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra', 7: 'hawama',8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi', 14: 'jabwar',15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan', 21:'noise'}\n\n\nscored_birds ={0: 'akiapo', 1: 'aniani', 2: 'apapan', 3: 'barpet', 4: 'crehon', 5: 'elepai', 6: 'ercfra',7:  'hawama', 8: 'hawcre', 9: 'hawgoo', 10: 'hawhaw', 11: 'hawpet1', 12: 'houfin', 13: 'iiwi',14: 'jabwar', 15: 'maupar', 16: 'omao', 17: 'puaioh', 18: 'skylar', 19: 'warwhe1', 20: 'yefcan'}\n\nprint(scored_birds)\n\nfile_list = [f.split('.')[0] for f in sorted(os.listdir(test_audio_dir))]\n\nprint('Number of test soundscapes:', len(file_list))\n\npred = {'row_id': [], 'target': []}\n\n# Process audio files and make predictions\nfor afile in file_list:\n    # read files\n    path = test_audio_dir + afile + '.ogg'\n    audio, sr = librosa.load(path)\n    audio = _resample_if_necessary(audio, sr)\n    audio = _mix_down_if_necessary(audio)\n\n    # chunk audio to 5sec\n    chunks = []\n    for i in range(0, len(audio), total_length):\n        sig = audio[i:i + total_length]\n        \n        length_signal = sig.shape[0]\n        num_missing_samples = total_length - length_signal\n        sig = np.concatenate([sig, [0] * num_missing_samples])\n        \n        mel = librosa.feature.melspectrogram(sig, hop_length=520,sr=target_sample_rate, \n        fmin=20,fmax=14_000, n_mels=128)\n        mel = librosa.amplitude_to_db(mel, top_db=100, ref=np.max)\n\n        mel = np.array(normalize(torch.tensor(mel).unsqueeze(0)))\n        mel = np.array(cv2.resize(mel.squeeze(), (64, 64))).reshape((1, 64, 64))\n        \n        # vgg11 need 3-channel image:\n        sig = np.zeros((3, 64, 64))\n        sig[0, :, :] = mel\n        sig[1, :, :] = mel\n        sig[2, :, :] = mel\n\n        chunks.append(torch.tensor(sig,dtype=torch.float))\n\n    test = torch.stack(chunks).to(device) (-1, 3, 64, 64)\n\n    with torch.no_grad():\n        outputs = best_model(test).detach().cpu().numpy().argmax(1)\n\n    for i in range(len(chunks)): \n        for bird in scored_birds.values():\n            chunk_end_time = (i + 1) * 5\n            target2index = dict_birds[outputs[i]]\n            target = bird == target2index\n            row_id = afile + '_' + bird + '_' + str(chunk_end_time)\n            pred['row_id'].append(row_id)\n            pred['target'].append(target)\n\npred = pd.DataFrame(pred)\n\nprint(pred) # all targets values\nprint()\nprint(pred[pred['target'] ==True]) # only True targets\n\npred.to_csv('submission.csv', index=False) \n```\n\nThe result i get when I run\n```\nNumber of test soundscapes: 1\n\n\n                              row_id  target\n0      soundscape_453028782_akiapo_5   False\n1      soundscape_453028782_aniani_5   False\n2      soundscape_453028782_apapan_5   False\n3      soundscape_453028782_barpet_5   False\n4      soundscape_453028782_crehon_5   False\n..                               ...     ...\n247     soundscape_453028782_omao_60   False\n248   soundscape_453028782_puaioh_60   False\n249   soundscape_453028782_skylar_60   False\n250  soundscape_453028782_warwhe1_60   False\n251   soundscape_453028782_yefcan_60   False\n[252 rows x 2 columns]\n\n                             row_id  target\n12    soundscape_453028782_houfin_5    True\n33   soundscape_453028782_houfin_10    True\n54   soundscape_453028782_houfin_15    True\n75   soundscape_453028782_houfin_20    True\n98   soundscape_453028782_jabwar_25    True\n117  soundscape_453028782_houfin_30    True\n138  soundscape_453028782_houfin_35    True\n159  soundscape_453028782_houfin_40    True\n180  soundscape_453028782_houfin_45    True\n201  soundscape_453028782_houfin_50    True\n224  soundscape_453028782_jabwar_55    True\n245  soundscape_453028782_jabwar_60    True\n```\n\nWhen I predict on the validation data set, I get A good f1-macro score but when I submit I don't get the same or nearly the same score!\n\nwhen I train I use  only the scored classes sounds:\nI train my model on augmented dataset  and my validation set is the original data set\n\nI think my problem is from the code when I submit but I don't know where.\nAny advice?",
    "1751110": "I'm actually struggling to get a correct submission too... It's been driving me mad for three days now ^^\n\nWhy do you set the sample_rate to 22kHz ? Shouldn't it be 32kHz ?",
    "1750616": "0.48 - all False\n0.51 - all True",
    "1751179": "Me too. I trained a good model .f1 =0.88. But only 0.49 on leaderboard?? I have predicted all pigeons (scoring and unscoring). Gaussian noise is also added .\nBut it's bad...",
    "1750622": "I know but I don't know why I get this while my model can predict different sounds.\nhere is my best score on my machine:\n```\nBEST EPOCH(5) time=(0.03)hours: loss=(train_loss=0.7492 | val_loss=0.677) | F1-score=(train_f1_score=0.568 | val_f1_score=0.61) | Accuracy=(train_accuracy=[0.77] | val_accuracy=0.793)\n\n```",
    "1751132": ""
  }
}