{
  "id": 497539,
  "title": "[LB0.68] my experimental results",
  "url": "/competitions/birdclef-2024/discussion/497539",
  "author_name": "lhwcv",
  "post_date": "2024-04-25T01:16:18.091000",
  "votes": 84,
  "comment_count": 109,
  "views": 0,
  "content": "<p>Basic Setting</p>\n<p>5fold,  Focal BCE, EMA</p>\n<p>model1:  2023 rank1 model<br>\nmodel2:  2023 rank4 model<br>\naug1:    norm to 0~255 then norm by imagenet RGB std,  then use Hflip, CoarseDropOut, Mixup<br>\naug2:   norm to 0~1,  one channel, use  Freq and Time aug,  Mixup</p>\n<p>LB:  0.63<br>\nmodel1 + aug2 + only first 5sec + with additional time domain noise aug <br>\nLB:  0.64<br>\nmodel1 + aug2 + only first 5sec<br>\nLB:  0.62<br>\nmodel1 + aug1 + only first 5sec  </p>\n<p>!! noise aug maybe harmful when data quality is low</p>\n<p>LB:  0.66  <br>\nmodel2 + aug1 + only first 5sec   (fold3 is  0.67)<br>\nLB:  0.6<br>\nmodel2 + aug1 + random crop 5sec when train</p>\n<p>!!  random crop may crop just noise, and overfit the noise<br>\nand I found that the data quality is very low, e.g  brfowl1  mix other bird sound</p>\n<p>First 5sec is good but not the best</p>\n<p>model2 all fold:</p>\n<p>fold1: CV: 0.974080, LB: 0.64<br>\nfold2: CV: 0.974512, LB: 0.66<br>\nfold3: CV: 0.971235, LB: 0.67<br>\nfold4: CV: 0.973694, LB: 0.64<br>\nfold5: CV: 0.970303, LB: 0.64</p>",
  "messages": [
    {
      "id": 2773933,
      "postDate": "2024-04-25T01:16:18.090Z",
      "content": "<p>Basic Setting</p>\n<p>5fold,  Focal BCE, EMA</p>\n<p>model1:  2023 rank1 model<br>\nmodel2:  2023 rank4 model<br>\naug1:    norm to 0~255 then norm by imagenet RGB std,  then use Hflip, CoarseDropOut, Mixup<br>\naug2:   norm to 0~1,  one channel, use  Freq and Time aug,  Mixup</p>\n<p>LB:  0.63<br>\nmodel1 + aug2 + only first 5sec + with additional time domain noise aug <br>\nLB:  0.64<br>\nmodel1 + aug2 + only first 5sec<br>\nLB:  0.62<br>\nmodel1 + aug1 + only first 5sec  </p>\n<p>!! noise aug maybe harmful when data quality is low</p>\n<p>LB:  0.66  <br>\nmodel2 + aug1 + only first 5sec   (fold3 is  0.67)<br>\nLB:  0.6<br>\nmodel2 + aug1 + random crop 5sec when train</p>\n<p>!!  random crop may crop just noise, and overfit the noise<br>\nand I found that the data quality is very low, e.g  brfowl1  mix other bird sound</p>\n<p>First 5sec is good but not the best</p>\n<p>model2 all fold:</p>\n<p>fold1: CV: 0.974080, LB: 0.64<br>\nfold2: CV: 0.974512, LB: 0.66<br>\nfold3: CV: 0.971235, LB: 0.67<br>\nfold4: CV: 0.973694, LB: 0.64<br>\nfold5: CV: 0.970303, LB: 0.64</p>",
      "rawMarkdown": "Basic Setting\n\n5fold,  Focal BCE, EMA\n\nmodel1:  2023 rank1 model\nmodel2:  2023 rank4 model\naug1:    norm to 0~255 then norm by imagenet RGB std,  then use Hflip, CoarseDropOut, Mixup\naug2:   norm to 0~1,  one channel, use  Freq and Time aug,  Mixup\n\nLB:  0.63\nmodel1 + aug2 + only first 5sec + with additional time domain noise aug \nLB:  0.64\nmodel1 + aug2 + only first 5sec\nLB:  0.62\nmodel1 + aug1 + only first 5sec  \n\n!! noise aug maybe harmful when data quality is low\n\nLB:  0.66  \nmodel2 + aug1 + only first 5sec   (fold3 is  0.67)\nLB:  0.6\nmodel2 + aug1 + random crop 5sec when train\n\n!!  random crop may crop just noise, and overfit the noise\nand I found that the data quality is very low, e.g  brfowl1  mix other bird sound\n\nFirst 5sec is good but not the best\n\nmodel2 all fold:\n\nfold1: CV: 0.974080, LB: 0.64\nfold2: CV: 0.974512, LB: 0.66\nfold3: CV: 0.971235, LB: 0.67\nfold4: CV: 0.973694, LB: 0.64\nfold5: CV: 0.970303, LB: 0.64",
      "votes": 84
    },
    {
      "id": 2783966,
      "postDate": "2024-04-30T03:24:08.210Z",
      "content": "<p>Update 04-30,  Here is my latest experiments (fig): <br>\nI tried pretrain,  add nocall,  train with extend class,  multi audio clip with score …<br>\nStill hard for me to find the correlation!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F2763710548b0f26f78580c547acaffaa%2F6951714446627_.pic.jpg?generation=1714447369360418&amp;alt=media\"></p>",
      "rawMarkdown": "Update 04-30,  Here is my latest experiments (fig): \nI tried pretrain,  add nocall,  train with extend class,  multi audio clip with score ...\nStill hard for me to find the correlation!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F2763710548b0f26f78580c547acaffaa%2F6951714446627_.pic.jpg?generation=1714447369360418&alt=media)",
      "votes": 9,
      "replies": [
        {
          "id": 2783976,
          "postDate": "2024-04-30T03:34:21.340Z",
          "content": "<p>Thank you for sharing! May I ask what does cfg0/1/2/3 means?</p>",
          "rawMarkdown": "Thank you for sharing! May I ask what does cfg0/1/2/3 means?",
          "replies": [
            {
              "id": 2783981,
              "postDate": "2024-04-30T03:47:37.687Z",
              "content": "<p>cfg0 is my baseline with label_smooth=0, second coef=1.0  and some augs<br>\ncfg1～cfg3 is diffrent second coef mainly (poor so I didn't submit all)</p>",
              "rawMarkdown": "cfg0 is my baseline with label_smooth=0, second coef=1.0  and some augs\ncfg1～cfg3 is diffrent second coef mainly (poor so I didn't submit all)",
              "votes": 2
            },
            {
              "id": 2783999,
              "postDate": "2024-04-30T04:07:55.583Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            }
          ]
        },
        {
          "id": 2785750,
          "postDate": "2024-05-01T02:43:53.180Z",
          "content": "<p>wow this is super impressive! may I ask which backbone are you using?</p>",
          "rawMarkdown": "wow this is super impressive! may I ask which backbone are you using?",
          "replies": [
            {
              "id": 2785758,
              "postDate": "2024-05-01T02:55:12.480Z",
              "content": "<p>I’m using eca_nfnet_l0</p>",
              "rawMarkdown": "I’m using eca_nfnet_l0",
              "votes": 5
            }
          ]
        },
        {
          "id": 2787737,
          "postDate": "2024-05-01T23:18:49.907Z",
          "content": "<p>Interesting that the pretraining did not improve the scores?! Did you able to figure it out? In my case, with 5 epochs pretraining with the previous comp data, it gives some boost.</p>",
          "rawMarkdown": "Interesting that the pretraining did not improve the scores?! Did you able to figure it out? In my case, with 5 epochs pretraining with the previous comp data, it gives some boost.",
          "votes": 1,
          "replies": [
            {
              "id": 2791956,
              "postDate": "2024-05-04T00:44:34.233Z",
              "content": "<p>I'm debugging this, strange for me it gives no boost. 🤣</p>",
              "rawMarkdown": "I'm debugging this, strange for me it gives no boost. 🤣"
            }
          ]
        }
      ]
    },
    {
      "id": 2780851,
      "postDate": "2024-04-28T13:03:20.260Z",
      "content": "<p>Well for the same experiment settings, I am unable to reproduce these results. Following 2023 4th place solution.<br>\nFirst 5 seconds training. </p>\n<p>Apparently, the only changes I can think of might be the followings: <br>\nEarly Stopping (Around 40 Epochs)? (As the validation set is not reliable, so small reduction in ROC or CMAP is not a good way to early stop.)<br>\nTime Mask / Freq Mask?<br>\neca_nfnet_l0 <br>\nOnnx inference<br>\nNormalizing the Mel Spec or Normalization of Wave?<br>\nMixed Precision?<br>\nGoogle Bird Model Predictions for KD (Version 2)</p>\n<p>Not sure, if I missed something. Remaining setting is same. (I am only able to get the score around 0.62 max)</p>",
      "rawMarkdown": "Well for the same experiment settings, I am unable to reproduce these results. Following 2023 4th place solution.\nFirst 5 seconds training. \n\nApparently, the only changes I can think of might be the followings: \nEarly Stopping (Around 40 Epochs)? (As the validation set is not reliable, so small reduction in ROC or CMAP is not a good way to early stop.)\nTime Mask / Freq Mask?\neca_nfnet_l0 \nOnnx inference\nNormalizing the Mel Spec or Normalization of Wave?\nMixed Precision?\nGoogle Bird Model Predictions for KD (Version 2)\n\nNot sure, if I missed something. Remaining setting is same. (I am only able to get the score around 0.62 max)",
      "votes": 7,
      "replies": [
        {
          "id": 2781086,
          "postDate": "2024-04-28T14:53:55.437Z",
          "content": "<p>Did you add the thresholding mentioned <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/497539#2774249\" target=\"_blank\">here</a>? I think this might be a key component.</p>",
          "rawMarkdown": "Did you add the thresholding mentioned [here](https://www.kaggle.com/competitions/birdclef-2024/discussion/497539#2774249)? I think this might be a key component.",
          "replies": [
            {
              "id": 2781116,
              "postDate": "2024-04-28T15:17:23.497Z",
              "content": "<p>good point i miss this details. i'm also struggling to reproduce results.</p>",
              "rawMarkdown": "good point i miss this details. i'm also struggling to reproduce results.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2781712,
          "postDate": "2024-04-29T00:44:37.827Z",
          "content": "<ol>\n<li>Time Mask / Freq Mask? --- Time and Freq</li>\n<li>Normalizing the Mel Spec or Normalization of Wave? -- On Mel spec, norm to 0~255 then normalize by imagenet RGB mean std.</li>\n<li>Mixed Precision? -- yes</li>\n</ol>\n<p>Based on the information you provided, I suggest you to check:</p>\n<p>Using FocalBCE<br>\nChecking sencond_coef, label_smooth</p>",
          "rawMarkdown": "1. Time Mask / Freq Mask? --- Time and Freq\n2. Normalizing the Mel Spec or Normalization of Wave? -- On Mel spec, norm to 0~255 then normalize by imagenet RGB mean std.\n3. Mixed Precision? -- yes\n\nBased on the information you provided, I suggest you to check:\n\nUsing FocalBCE\nChecking sencond_coef, label_smooth",
          "replies": [
            {
              "id": 2781727,
              "postDate": "2024-04-29T01:03:02.267Z",
              "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a>   and no KD (this slightly decreased my cv), sharing more detailed information about your CV might be beneficial for analysis.  my CV range from  0.972 -- 0.974 </p>",
              "rawMarkdown": "@salmanahmedtamu   and no KD (this slightly decreased my cv), sharing more detailed information about your CV might be beneficial for analysis.  my CV range from  0.972 -- 0.974 "
            },
            {
              "id": 2781823,
              "postDate": "2024-04-29T02:56:48.237Z",
              "content": "<p>Yeah, I meant did you use Time and Freq Augs?</p>\n<p>My CV ranges from 0.98 to 0.99</p>\n<p>I used label smooth to 0.05, and secondary_coefs to 1.0</p>\n<p>0.66 or 0.67 is without KD? </p>",
              "rawMarkdown": "Yeah, I meant did you use Time and Freq Augs?\n\nMy CV ranges from 0.98 to 0.99\n\nI used label smooth to 0.05, and secondary_coefs to 1.0\n\n0.66 or 0.67 is without KD? \n"
            },
            {
              "id": 2781825,
              "postDate": "2024-04-29T02:58:23.810Z",
              "content": "<p>yes I used.  without KD and without pretrain,use imagenet pretrained,  0.98 to 0.99 is very high, I guess your validation don't use secondary_coef 1.0 ?</p>",
              "rawMarkdown": "yes I used.  without KD and without pretrain,use imagenet pretrained,  0.98 to 0.99 is very high, I guess your validation don't use secondary_coef 1.0 ?"
            },
            {
              "id": 2781829,
              "postDate": "2024-04-29T03:02:44.960Z",
              "content": "<p>and my fold2 log:<br>\nsteps: 395,epo: 1, best_score: 0.923567 train_loss: 0.029568 , eval_loss: 0.023104946649113257 , roc_auc: 0.9235667694966543 <br>\nsteps: 790,epo: 2, best_score: 0.958245 train_loss: 0.018857 , eval_loss: 0.017212639977042744 , roc_auc: 0.9582446547051086 <br>\nsteps: 1185,epo: 3, best_score: 0.966371 train_loss: 0.016296 , eval_loss: 0.015339857312040283 , roc_auc: 0.9663713526370147 <br>\nsteps: 1580,epo: 4, best_score: 0.969536 train_loss: 0.014625 , eval_loss: 0.014284247196847935 , roc_auc: 0.9695358129988351 <br>\nsteps: 1975,epo: 5, best_score: 0.971608 train_loss: 0.013271 , eval_loss: 0.013741090936078266 , roc_auc: 0.9716079793856934 <br>\nsteps: 2370,epo: 6, best_score: 0.974090 train_loss: 0.010935 , eval_loss: 0.013370304558934136 , roc_auc: 0.9740900915645582 <br>\nsteps: 2765,epo: 7, best_score: 0.974090 train_loss: 0.009982 , eval_loss: 0.013458995466035876 , roc_auc: 0.9736567230864034 <br>\nsteps: 3160,epo: 8, best_score: 0.974512 train_loss: 0.009245 , eval_loss: 0.013592320273618227 , roc_auc: 0.9745117067793344 <br>\nsteps: 3555,epo: 9, best_score: 0.974512 train_loss: 0.008837 , eval_loss: 0.013835379566586056 , roc_auc: 0.9740697744987592 <br>\nsteps: 3950,epo: 10, best_score: 0.974512 train_loss: 0.008407 , eval_loss: 0.014064931419569177 , roc_auc: 0.9725833905498832 <br>\nsteps: 4345,epo: 11, best_score: 0.974512 train_loss: 0.007450 , eval_loss: 0.015370996599388006 , roc_auc: 0.9704292221454025 <br>\nsteps: 4740,epo: 12, best_score: 0.974512 train_loss: 0.006486 , eval_loss: 0.016247244984177605 , roc_auc: 0.9698166650036316 <br>\nsteps: 5135,epo: 13, best_score: 0.974512 train_loss: 0.006573 , eval_loss: 0.016679136829402346 , roc_auc: 0.9687894194276919 <br>\n2024-04-15_18:13:52: early stopped!</p>",
              "rawMarkdown": "and my fold2 log:\nsteps: 395,epo: 1, best_score: 0.923567 train_loss: 0.029568 , eval_loss: 0.023104946649113257 , roc_auc: 0.9235667694966543 \nsteps: 790,epo: 2, best_score: 0.958245 train_loss: 0.018857 , eval_loss: 0.017212639977042744 , roc_auc: 0.9582446547051086 \nsteps: 1185,epo: 3, best_score: 0.966371 train_loss: 0.016296 , eval_loss: 0.015339857312040283 , roc_auc: 0.9663713526370147 \nsteps: 1580,epo: 4, best_score: 0.969536 train_loss: 0.014625 , eval_loss: 0.014284247196847935 , roc_auc: 0.9695358129988351 \nsteps: 1975,epo: 5, best_score: 0.971608 train_loss: 0.013271 , eval_loss: 0.013741090936078266 , roc_auc: 0.9716079793856934 \nsteps: 2370,epo: 6, best_score: 0.974090 train_loss: 0.010935 , eval_loss: 0.013370304558934136 , roc_auc: 0.9740900915645582 \nsteps: 2765,epo: 7, best_score: 0.974090 train_loss: 0.009982 , eval_loss: 0.013458995466035876 , roc_auc: 0.9736567230864034 \nsteps: 3160,epo: 8, best_score: 0.974512 train_loss: 0.009245 , eval_loss: 0.013592320273618227 , roc_auc: 0.9745117067793344 \nsteps: 3555,epo: 9, best_score: 0.974512 train_loss: 0.008837 , eval_loss: 0.013835379566586056 , roc_auc: 0.9740697744987592 \nsteps: 3950,epo: 10, best_score: 0.974512 train_loss: 0.008407 , eval_loss: 0.014064931419569177 , roc_auc: 0.9725833905498832 \nsteps: 4345,epo: 11, best_score: 0.974512 train_loss: 0.007450 , eval_loss: 0.015370996599388006 , roc_auc: 0.9704292221454025 \nsteps: 4740,epo: 12, best_score: 0.974512 train_loss: 0.006486 , eval_loss: 0.016247244984177605 , roc_auc: 0.9698166650036316 \nsteps: 5135,epo: 13, best_score: 0.974512 train_loss: 0.006573 , eval_loss: 0.016679136829402346 , roc_auc: 0.9687894194276919 \n2024-04-15_18:13:52: early stopped!"
            },
            {
              "id": 2781830,
              "postDate": "2024-04-29T03:03:17.320Z",
              "content": "<p>Yeah, I don't have secondary_coefs in validation.</p>",
              "rawMarkdown": "Yeah, I don't have secondary_coefs in validation."
            },
            {
              "id": 2781834,
              "postDate": "2024-04-29T03:08:21.923Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2781838,
              "postDate": "2024-04-29T03:15:28.583Z",
              "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Do you use sampling, or weights for the loss?</p>",
              "rawMarkdown": "@lihaoweicvch Do you use sampling, or weights for the loss?"
            },
            {
              "id": 2781846,
              "postDate": "2024-04-29T03:18:05.493Z",
              "content": "<p>currently no, but I'll add this later, sampling and weighting by rating</p>",
              "rawMarkdown": "currently no, but I'll add this later, sampling and weighting by rating"
            },
            {
              "id": 2782114,
              "postDate": "2024-04-29T06:17:17.897Z",
              "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> I added secondary labels in validation and used focalbce. <br>\nNow my CV is 0.9575</p>",
              "rawMarkdown": "@lihaoweicvch I added secondary labels in validation and used focalbce. \nNow my CV is 0.9575",
              "votes": 1
            },
            {
              "id": 2783700,
              "postDate": "2024-04-29T22:13:29.320Z",
              "content": "<p>what is secondary_coef? are you using the secondary label as birds to detect ?</p>",
              "rawMarkdown": "what is secondary_coef? are you using the secondary label as birds to detect ?"
            },
            {
              "id": 2783756,
              "postDate": "2024-04-29T23:52:02.460Z",
              "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> It means, target value for secondary labels, if its going to be 1 or 0.8 or whatever you want to set. </p>",
              "rawMarkdown": "@ludovick It means, target value for secondary labels, if its going to be 1 or 0.8 or whatever you want to set. ",
              "votes": 1
            },
            {
              "id": 2786180,
              "postDate": "2024-05-01T07:48:40.137Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2773935,
      "postDate": "2024-04-25T01:19:09.330Z",
      "content": "<p>Now my focus work is to explore the influence of the data quality, follows are doing:</p>\n<ul>\n<li>use trained model select out the high quality segs by predicted score, use Google Kaggle Bird Model maybe be a good choice?</li>\n</ul>\n<p>According to the training data prediction of Google Model in 2024, if the first 5 seconds of the clip are selected, the metric is at 0.97325. If selected randomly, it is between 0.94-0.95. However, if the segment with the highest category score is chosen, it goes up to 0.983.</p>",
      "rawMarkdown": "Now my focus work is to explore the influence of the data quality, follows are doing:\n- use trained model select out the high quality segs by predicted score, use Google Kaggle Bird Model maybe be a good choice?\n\nAccording to the training data prediction of Google Model in 2024, if the first 5 seconds of the clip are selected, the metric is at 0.97325. If selected randomly, it is between 0.94-0.95. However, if the segment with the highest category score is chosen, it goes up to 0.983.\n\n",
      "votes": 7
    },
    {
      "id": 2781790,
      "postDate": "2024-04-29T02:31:59.587Z",
      "content": "<p>Here is my latest experiments (fig): <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F1dbd036d6cab4700009843ef7503dbea%2F6911714357731_.pic.jpg?generation=1714357891979872&amp;alt=media\"></p>",
      "rawMarkdown": "Here is my latest experiments (fig): \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F1dbd036d6cab4700009843ef7503dbea%2F6911714357731_.pic.jpg?generation=1714357891979872&alt=media)",
      "votes": 3
    },
    {
      "id": 2776589,
      "postDate": "2024-04-26T08:40:49.267Z",
      "content": "<p>Has anyone achieved good results with ViT? （figure below is  my results using different backbone）In addition, in this competition, we need to balance cost-effectiveness. It is possible that an ensemble of smaller models may win.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F35c834f2a37f6f3fd44c07b1e3c1d37f%2F6891714120735_.pic.jpg?generation=1714120847514341&amp;alt=media\"></p>",
      "rawMarkdown": "Has anyone achieved good results with ViT? （figure below is  my results using different backbone）In addition, in this competition, we need to balance cost-effectiveness. It is possible that an ensemble of smaller models may win.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F35c834f2a37f6f3fd44c07b1e3c1d37f%2F6891714120735_.pic.jpg?generation=1714120847514341&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 2776790,
          "postDate": "2024-04-26T10:37:51.300Z",
          "content": "<p>efficientvit_b2  has export Bug, the LB score is not correct, I'll update later</p>",
          "rawMarkdown": "efficientvit_b2  has export Bug, the LB score is not correct, I'll update later",
          "votes": 2
        },
        {
          "id": 2777069,
          "postDate": "2024-04-26T13:55:37.833Z",
          "content": "<p>I started with <code>efficientvit</code>, but moved to <code>efficientnet</code> as the models can be converted to onnx without modifying the padding layers. Also, in my early experiments, CV/LB was stronger with <code>efficientnet</code>.</p>",
          "rawMarkdown": "I started with `efficientvit`, but moved to `efficientnet` as the models can be converted to onnx without modifying the padding layers. Also, in my early experiments, CV/LB was stronger with `efficientnet`.",
          "votes": 2,
          "replies": [
            {
              "id": 2778162,
              "postDate": "2024-04-27T02:07:00.153Z",
              "content": "<p>After fixed the bug, my efficienvit got 0.65 which is also slightly lower than efficientnet.</p>",
              "rawMarkdown": "After fixed the bug, my efficienvit got 0.65 which is also slightly lower than efficientnet.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2780018,
          "postDate": "2024-04-28T01:41:08.017Z",
          "content": "<p>what is the meaning of LB？</p>",
          "rawMarkdown": "what is the meaning of LB？",
          "votes": 1,
          "replies": [
            {
              "id": 2780049,
              "postDate": "2024-04-28T02:32:34.780Z",
              "content": "<p>Leaderboard,  Public  </p>",
              "rawMarkdown": "Leaderboard,  Public  ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2776080,
      "postDate": "2024-04-26T01:20:04.860Z",
      "content": "<p>I have a question: in the online test set, would a 5-second segment include multiple birds? If there are multiple birds, what coefficient would be used? Is it 1?</p>",
      "rawMarkdown": "I have a question: in the online test set, would a 5-second segment include multiple birds? If there are multiple birds, what coefficient would be used? Is it 1?",
      "votes": 4,
      "replies": [
        {
          "id": 2777049,
          "postDate": "2024-04-26T13:37:32.190Z",
          "content": "<p>Yes, there can be multiple birds, yes all birds that are audible in the 5-second segment have the ground truth 1</p>",
          "rawMarkdown": "Yes, there can be multiple birds, yes all birds that are audible in the 5-second segment have the ground truth 1",
          "votes": 2,
          "replies": [
            {
              "id": 2778164,
              "postDate": "2024-04-27T02:07:46.780Z",
              "content": "<p>Thank you for sharing </p>",
              "rawMarkdown": "Thank you for sharing ",
              "votes": 1
            },
            {
              "id": 2784770,
              "postDate": "2024-04-30T13:37:45.667Z",
              "content": "<p>Does that mean a multiclass approach is a dead end?</p>",
              "rawMarkdown": "Does that mean a multiclass approach is a dead end?"
            },
            {
              "id": 2787866,
              "postDate": "2024-05-02T01:14:39.230Z",
              "content": "<blockquote>\n  <p>all birds that are audible in the 5-second segment have the ground truth 1</p>\n</blockquote>\n<p>What makes you think so?</p>",
              "rawMarkdown": "> all birds that are audible in the 5-second segment have the ground truth 1\n\nWhat makes you think so?",
              "votes": 1
            },
            {
              "id": 2787964,
              "postDate": "2024-05-02T03:31:23.673Z",
              "content": "<p>Because that describes a multilabel problem.</p>",
              "rawMarkdown": "Because that describes a multilabel problem.",
              "votes": 1
            },
            {
              "id": 2789352,
              "postDate": "2024-05-02T16:41:56.643Z",
              "content": "<p>I am also struggling to reach my same LB performance after using FocalBCE instead of CrossEntropy though.</p>",
              "rawMarkdown": "I am also struggling to reach my same LB performance after using FocalBCE instead of CrossEntropy though."
            }
          ]
        }
      ]
    },
    {
      "id": 2774228,
      "postDate": "2024-04-25T05:11:20.933Z",
      "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>  <a href=\"https://www.kaggle.com/lmhongkhnh\" target=\"_blank\">@lmhongkhnh</a>  I didn't find correlation, even though CV got  0.985 (with KD), LB may get  0.64<br>\nfold1:  CV: 0.974080, LB:  0.64<br>\nfold2:  CV: 0.974512, LB:  0.66<br>\nfold3:  CV: 0.971235, LB:  0.67<br>\nfold4:  CV: 0.973694, LB:  0.64<br>\nfold5: CV: 0.970303, LB:  0.64</p>",
      "rawMarkdown": "@cody11null  @lmhongkhnh  I didn't find correlation, even though CV got  0.985 (with KD), LB may get  0.64\nfold1:  CV: 0.974080, LB:  0.64\nfold2:  CV: 0.974512, LB:  0.66\nfold3:  CV: 0.971235, LB:  0.67\nfold4:  CV: 0.973694, LB:  0.64\nfold5: CV: 0.970303, LB:  0.64\n",
      "votes": 4,
      "replies": [
        {
          "id": 2775234,
          "postDate": "2024-04-25T14:56:46.343Z",
          "content": "<p>How did you setup your CV? Only on primary labels? And did you evaluate on samples &gt;= than a specific quality rating? Or on all data. </p>",
          "rawMarkdown": "How did you setup your CV? Only on primary labels? And did you evaluate on samples >= than a specific quality rating? Or on all data. ",
          "votes": 1,
          "replies": [
            {
              "id": 2775998,
              "postDate": "2024-04-26T00:11:41.990Z",
              "content": "<p>I used secondary_coef = 1.0 and evaluate on all without specific quality rating.<br>\nHowever, I don't think this is the most reasonable approach, we need more experiments, especially when there are multiple tags. Taking only 5 seconds is obviously not necessarily enough to capture multiple birds calling.</p>\n<pre><code> s  secondary_labels:\n       s !=  and s  CFG():\n             target] = self\n</code></pre>",
              "rawMarkdown": "I used secondary_coef = 1.0 and evaluate on all without specific quality rating.\nHowever, I don't think this is the most reasonable approach, we need more experiments, especially when there are multiple tags. Taking only 5 seconds is obviously not necessarily enough to capture multiple birds calling.\n```\nfor s in secondary_labels:\n      if s != \"\" and s in CFG.bird2id.keys():\n             target[CFG.bird2id[s]] = self.secondary_coef\n```",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2799893,
      "postDate": "2024-05-08T02:53:01.450Z",
      "content": "<p>Nice work! May I ask how you ensemble? When I try to use concurrent.futures.ThreadPoolExecutor to perform inference on multiple models simultaneously, the execution runs smoothly but errors occur during submission.</p>",
      "rawMarkdown": "Nice work! May I ask how you ensemble? When I try to use concurrent.futures.ThreadPoolExecutor to perform inference on multiple models simultaneously, the execution runs smoothly but errors occur during submission.",
      "votes": 1,
      "replies": [
        {
          "id": 2802272,
          "postDate": "2024-05-09T00:41:45.063Z",
          "content": "<p>try batch_size=48  with onnx session option num_threads=4</p>",
          "rawMarkdown": "try batch_size=48  with onnx session option num_threads=4",
          "votes": 2,
          "replies": [
            {
              "id": 2803250,
              "postDate": "2024-05-09T11:58:39.563Z",
              "content": "<p>My onnx is slower than jit with batch_size=4, with batch_size=48 it could be faster?</p>",
              "rawMarkdown": "My onnx is slower than jit with batch_size=4, with batch_size=48 it could be faster?",
              "votes": 2
            },
            {
              "id": 2803406,
              "postDate": "2024-05-09T13:42:13.290Z",
              "content": "<p>I don't know why all my submissions using num_threads=4 have failed….</p>",
              "rawMarkdown": "I don't know why all my submissions using num_threads=4 have failed....",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2795816,
      "postDate": "2024-05-06T02:19:32.993Z",
      "content": "<p>Have you tried efficientnet models ? How well do they perform for you ? I tried comparing few models like efficientnet, mobile net and efficientnet v2, there isn't much significant change when I performed them on the same pipeline,  what's your results on this ? </p>",
      "rawMarkdown": "Have you tried efficientnet models ? How well do they perform for you ? I tried comparing few models like efficientnet, mobile net and efficientnet v2, there isn't much significant change when I performed them on the same pipeline,  what's your results on this ? ",
      "votes": 1,
      "replies": [
        {
          "id": 2796122,
          "postDate": "2024-05-06T06:04:21.510Z",
          "content": "<p>I tried efficientnet b0 to b2 and eca_nfnet_l0,   different folds perform differently, there's no absolute best across different  backbone, I usually fluctuate between 0.64--0.68. Your situation seems nice, pretty stable, could you share more about your data pipeline if you don't mind?</p>",
          "rawMarkdown": "I tried efficientnet b0 to b2 and eca_nfnet_l0,   different folds perform differently, there's no absolute best across different  backbone, I usually fluctuate between 0.64--0.68. Your situation seems nice, pretty stable, could you share more about your data pipeline if you don't mind?",
          "votes": 2,
          "replies": [
            {
              "id": 2796222,
              "postDate": "2024-05-06T06:43:17.260Z",
              "content": "<p>tbh, my situation isn't stable either, I cannot recreate my best scoring model myself, a lot of variables and randomness is involved, and i am not using the folds strategy, the cv isn't so reliable anyway so I am using most of the data chunk with a little bit of cleaning  in the training process</p>",
              "rawMarkdown": "tbh, my situation isn't stable either, I cannot recreate my best scoring model myself, a lot of variables and randomness is involved, and i am not using the folds strategy, the cv isn't so reliable anyway so I am using most of the data chunk with a little bit of cleaning  in the training process",
              "votes": 3
            },
            {
              "id": 2796737,
              "postDate": "2024-05-06T11:39:16.543Z",
              "content": "<p>Thank you for sharing， it is the same case on my side， LB is very sensitive </p>",
              "rawMarkdown": "Thank you for sharing， it is the same case on my side， LB is very sensitive ",
              "votes": 3
            },
            {
              "id": 2803251,
              "postDate": "2024-05-09T11:59:38.333Z",
              "content": "<p>It seems that same CV for different models do not correlate with LB score?</p>",
              "rawMarkdown": "It seems that same CV for different models do not correlate with LB score?",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2782100,
      "postDate": "2024-04-29T06:00:27.380Z",
      "content": "<p>Thank you for sharing your solution. Could you please clarify how you transform the waveform into a MelSpectrogram, whether it's done online or offline? Additionally, the 4th place solution in 2023 transforms the waveform into a MelSpectrogram during the forward propagation, for instance, e.g., </p>\n<pre><code>     ():\n        x = []\n         CFG.use_pcen:\n            x = self.pcen(x).unsqueeze()\n            self.pcen.reset()\n        :\n            x = self.logmelspec_extractor(x)[:, ] \n\n        \n        \n        \n        \n        \n        \n</code></pre>\n<p>How can I incorporate data transformation using Albumentations in the model's forward pass? Thanks a lot.</p>",
      "rawMarkdown": "Thank you for sharing your solution. Could you please clarify how you transform the waveform into a MelSpectrogram, whether it's done online or offline? Additionally, the 4th place solution in 2023 transforms the waveform into a MelSpectrogram during the forward propagation, for instance, e.g., \n```python\n    def forward(self, input):\n        x = input['wave']\n        if CFG.use_pcen:\n            x = self.pcen(x).unsqueeze(1)\n            self.pcen.reset()\n        else:\n            x = self.logmelspec_extractor(x)[:, None] # (32, 1, 128, 313)\n        \n        # TODO\n        # (32, 1, 128, 313) -> (32, 3, 128, 313)\n        # norm by imagenet RGB std\n        # then use Hflip\n        # CoarseDropOut\n        # Mixup\n```\n\nHow can I incorporate data transformation using Albumentations in the model's forward pass? Thanks a lot.",
      "votes": 1,
      "replies": [
        {
          "id": 2782423,
          "postDate": "2024-04-29T09:52:29.747Z",
          "content": "<p>you can try librosa.melspec in dataset</p>",
          "rawMarkdown": "you can try librosa.melspec in dataset",
          "votes": 1
        }
      ]
    },
    {
      "id": 2781897,
      "postDate": "2024-04-29T04:06:15.403Z",
      "content": "<p>nicee work </p>",
      "rawMarkdown": "nicee work ",
      "votes": 1
    },
    {
      "id": 2781863,
      "postDate": "2024-04-29T03:25:26.037Z",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Imagenet mean and std are something like this. <br>\nnormalize = transforms.Normalize(mean=[0.485, 0.456, 0.406],<br>\n                                 std=[0.229, 0.224, 0.225])</p>\n<p>I have 2 assumptions here.</p>\n<ol>\n<li><p>you multiplied these mean and std by 255.0 and then normalized. <br>\nWhich means you input to the model might not be in range of 0 - 1</p></li>\n<li><p>You normalize to 0-1 and then applied this imagenet norm.</p></li>\n</ol>",
      "rawMarkdown": "@lihaoweicvch Imagenet mean and std are something like this. \nnormalize = transforms.Normalize(mean=[0.485, 0.456, 0.406],\n                                 std=[0.229, 0.224, 0.225])\n\nI have 2 assumptions here.\n1. you multiplied these mean and std by 255.0 and then normalized. \nWhich means you input to the model might not be in range of 0 - 1\n\n2. You normalize to 0-1 and then applied this imagenet norm.",
      "votes": 1,
      "replies": [
        {
          "id": 2781870,
          "postDate": "2024-04-29T03:29:25.633Z",
          "content": "<p>yes you are right,  like this, input image is normalize to 0~255 first</p>\n<pre><code>def normalize(img, , , max_pixel_value=):\n     = .(, dtype=.float32)\n     *= max_pixel_value\n\n     = .(, dtype=.float32)\n     *= max_pixel_value\n\n    denominator = .reciprocal(, dtype=.float32)\n\n    img = img.astype(.float32)\n    img -= \n    img *= denominator\n     img\n</code></pre>",
          "rawMarkdown": "yes you are right,  like this, input image is normalize to 0~255 first\n```\ndef normalize(img, mean, std, max_pixel_value=255.0):\n    mean = np.array(mean, dtype=np.float32)\n    mean *= max_pixel_value\n\n    std = np.array(std, dtype=np.float32)\n    std *= max_pixel_value\n\n    denominator = np.reciprocal(std, dtype=np.float32)\n\n    img = img.astype(np.float32)\n    img -= mean\n    img *= denominator\n    return img\n```",
          "votes": 1,
          "replies": [
            {
              "id": 2787565,
              "postDate": "2024-05-01T20:06:54.650Z",
              "content": "<p>I'm a bit confused on what you do here, can't you just use np.interp to get the values of the image directly to [0,1] and then do the imagenet normalization? Also one curiosity, do you use wighted sampling for your dataloader?</p>",
              "rawMarkdown": "I'm a bit confused on what you do here, can't you just use np.interp to get the values of the image directly to [0,1] and then do the imagenet normalization? Also one curiosity, do you use wighted sampling for your dataloader?",
              "votes": 1
            },
            {
              "id": 2791959,
              "postDate": "2024-05-04T00:47:27.397Z",
              "content": "<p>weighted sampling  low down my score, see my latest experiments </p>",
              "rawMarkdown": "weighted sampling  low down my score, see my latest experiments ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2781767,
      "postDate": "2024-04-29T02:15:21.593Z",
      "content": "<p>Here is a question: If we have trained with more data segments and have observed significant improvement locally, would the recall rate usually increase when applied to another dataset?</p>",
      "rawMarkdown": "Here is a question: If we have trained with more data segments and have observed significant improvement locally, would the recall rate usually increase when applied to another dataset?",
      "votes": 1
    },
    {
      "id": 2781662,
      "postDate": "2024-04-28T23:11:20.323Z",
      "content": "<p>Thank you for posting your findings!</p>\n<blockquote>\n  <p>5fold, Focal BCE, EMA</p>\n</blockquote>\n<p>What is EMA?</p>\n<blockquote>\n  <p>LB: 0.63</p>\n</blockquote>\n<p>How do you compute the LB score for a model that had 5-fold CV? Do you submit all 5 models then average their LBs, or create 1 submission from an ensemble of all 5 models, or just submit 1 of the 5 that performed the best?</p>",
      "rawMarkdown": "Thank you for posting your findings!\n\n> 5fold, Focal BCE, EMA\n\nWhat is EMA?\n\n> LB: 0.63\n\nHow do you compute the LB score for a model that had 5-fold CV? Do you submit all 5 models then average their LBs, or create 1 submission from an ensemble of all 5 models, or just submit 1 of the 5 that performed the best?",
      "votes": 1,
      "replies": [
        {
          "id": 2781706,
          "postDate": "2024-04-29T00:39:57.707Z",
          "content": "<p>EMA： Exponential Moving Average</p>",
          "rawMarkdown": "EMA： Exponential Moving Average",
          "votes": 3,
          "replies": [
            {
              "id": 2781709,
              "postDate": "2024-04-29T00:41:24.633Z",
              "content": "<p>I submit 5 folds with model2 (LB  0.64~0.67),    model1  I only  submit fold2</p>",
              "rawMarkdown": "I submit 5 folds with model2 (LB  0.64~0.67),    model1  I only  submit fold2",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2781614,
      "postDate": "2024-04-28T21:25:49.663Z",
      "content": "<p>Should we try to use the centre 5 seconds of the audio clipping rather than the first 5 seconds. It is just an intuitive thought. Will this work or is there something fundamentally wrong with it ??</p>",
      "rawMarkdown": "Should we try to use the centre 5 seconds of the audio clipping rather than the first 5 seconds. It is just an intuitive thought. Will this work or is there something fundamentally wrong with it ??",
      "votes": 1,
      "replies": [
        {
          "id": 2781648,
          "postDate": "2024-04-28T23:00:28.100Z",
          "content": "<p>We want to avoid using 5 second clips that have no birdcall in our train data but are labeled as having some call. </p>\n<p>If a recording in the train set is 15 seconds and the label is pigeon it means somewhere in the 15 seconds there was a pigeon call. If we think about how the data was created it seems likely that this 15 second audio clip was taken from a longer recording at some location. When clipping the audio it makes sense that the start and end of the clip would be made near a pigeon call (why would you clip a recording to start 30 seconds before the first birdcall or end 30 seconds after any birdcall). Therefore first 5 seconds and last 5 seconds of any train data seem to be logical choices for less noisy data (but still possible that they have no birdcall).</p>",
          "rawMarkdown": "We want to avoid using 5 second clips that have no birdcall in our train data but are labeled as having some call. \n\nIf a recording in the train set is 15 seconds and the label is pigeon it means somewhere in the 15 seconds there was a pigeon call. If we think about how the data was created it seems likely that this 15 second audio clip was taken from a longer recording at some location. When clipping the audio it makes sense that the start and end of the clip would be made near a pigeon call (why would you clip a recording to start 30 seconds before the first birdcall or end 30 seconds after any birdcall). Therefore first 5 seconds and last 5 seconds of any train data seem to be logical choices for less noisy data (but still possible that they have no birdcall).",
          "votes": 1
        },
        {
          "id": 2781711,
          "postDate": "2024-04-29T00:44:04.953Z",
          "content": "<p>you can try center 5 seconds and maybe share to us, but I think  finding the audio clip which really contain bird calls can be more important.</p>",
          "rawMarkdown": "you can try center 5 seconds and maybe share to us, but I think  finding the audio clip which really contain bird calls can be more important.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2780836,
      "postDate": "2024-04-28T12:56:58.590Z",
      "content": "<p>Thanks for sharing again! Do you train only with the comp data and for how many epochs are you training for?</p>",
      "rawMarkdown": "Thanks for sharing again! Do you train only with the comp data and for how many epochs are you training for?",
      "votes": 1,
      "replies": [
        {
          "id": 2781713,
          "postDate": "2024-04-29T00:48:35.010Z",
          "content": "<p>Yes now I only use 2024 comp data,   train with EarlyStopping, usually  7-10 epoch with get the best </p>",
          "rawMarkdown": "Yes now I only use 2024 comp data,   train with EarlyStopping, usually  7-10 epoch with get the best ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2780830,
      "postDate": "2024-04-28T12:55:19.640Z",
      "content": "<p>How many epochs have you been training for?</p>",
      "rawMarkdown": "How many epochs have you been training for?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2781716,
          "postDate": "2024-04-29T00:50:58.280Z",
          "content": "<p>I train max 50 epochs, but train with EarlyStopping, usually 7-10 epoch with get the best, and I always submit the best</p>",
          "rawMarkdown": "I train max 50 epochs, but train with EarlyStopping, usually 7-10 epoch with get the best, and I always submit the best",
          "votes": 2,
          "replies": [
            {
              "id": 2796311,
              "postDate": "2024-05-06T07:18:18.827Z",
              "content": "<p>How do you determine the best? (I’m noticing that because local validation is unreliable, I don’t know when to stop training.)</p>",
              "rawMarkdown": "How do you determine the best? (I’m noticing that because local validation is unreliable, I don’t know when to stop training.)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2780069,
      "postDate": "2024-04-28T03:02:35.693Z",
      "content": "<p>Using multiple CVs to examine the correlation with LB may make it easier to discover the correlation ？</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0cb83b7c376ff5c13ae0b637dbe6a40a%2F6901714272397_.pic.jpg?generation=1714273338317490&amp;alt=media\"></p>",
      "rawMarkdown": "Using multiple CVs to examine the correlation with LB may make it easier to discover the correlation ？\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0cb83b7c376ff5c13ae0b637dbe6a40a%2F6901714272397_.pic.jpg?generation=1714273338317490&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 2781074,
          "postDate": "2024-04-28T14:44:31.673Z",
          "content": "<p>question  about your training, about the model , are you applying also the Kownledge Distillation and augmentation mentionned in the 4th solution? or are you just using the same backbone?</p>",
          "rawMarkdown": "question  about your training, about the model , are you applying also the Kownledge Distillation and augmentation mentionned in the 4th solution? or are you just using the same backbone?",
          "votes": 1,
          "replies": [
            {
              "id": 2781718,
              "postDate": "2024-04-29T00:51:56.673Z",
              "content": "<p>I  only use the same arch and backbone, no KD now</p>",
              "rawMarkdown": "I  only use the same arch and backbone, no KD now",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2777565,
      "postDate": "2024-04-26T17:54:45.460Z",
      "content": "<p>Could you tell me the score change when you used your notebook Bird Sound Denoise by Deep Model for train data and test data? And I wonder if you have tested 7 seconds, 10 seconds, etc. instead of the first 5 seconds.</p>",
      "rawMarkdown": "Could you tell me the score change when you used your notebook Bird Sound Denoise by Deep Model for train data and test data? And I wonder if you have tested 7 seconds, 10 seconds, etc. instead of the first 5 seconds.",
      "votes": 1,
      "replies": [
        {
          "id": 2778167,
          "postDate": "2024-04-27T02:09:28.177Z",
          "content": "<p>I didn’t use denoise now because that model hurt the bird sound. 7sec 10sec you mean train or infer ? </p>",
          "rawMarkdown": "I didn’t use denoise now because that model hurt the bird sound. 7sec 10sec you mean train or infer ? ",
          "votes": 1,
          "replies": [
            {
              "id": 2778567,
              "postDate": "2024-04-27T08:33:01.063Z",
              "content": "<p>There will be damage to the sound, but I'm quite curious about the result. I was talking about the case of training. I didn't see the increase in inference well, but I wondered what would happen to inference as well.</p>",
              "rawMarkdown": "There will be damage to the sound, but I'm quite curious about the result. I was talking about the case of training. I didn't see the increase in inference well, but I wondered what would happen to inference as well.",
              "votes": 1
            },
            {
              "id": 2778575,
              "postDate": "2024-04-27T08:38:12.120Z",
              "content": "<p>I tried 10s train in the early stage, then infer with 5s LB 0.64,  infer with 10s repeating to 5s LB 0.6</p>",
              "rawMarkdown": "I tried 10s train in the early stage, then infer with 5s LB 0.64,  infer with 10s repeating to 5s LB 0.6",
              "votes": 1
            },
            {
              "id": 2778679,
              "postDate": "2024-04-27T09:22:58.593Z",
              "content": "<p>Thank you for your kind reply!</p>",
              "rawMarkdown": "Thank you for your kind reply!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2775698,
      "postDate": "2024-04-25T18:42:03.163Z",
      "content": "<p>interesting, thanks for sharing;</p>\n<p>what kind of augmentation are you using regarding \"use Freq and Time aug, \". i try to use somes but haven't see big diff</p>",
      "rawMarkdown": "interesting, thanks for sharing;\n\nwhat kind of augmentation are you using regarding \"use Freq and Time aug, \". i try to use somes but haven't see big diff",
      "votes": 1,
      "replies": [
        {
          "id": 2775723,
          "postDate": "2024-04-25T18:52:21.220Z",
          "content": "<p>They are time and frequency domain masking. On albumentation it's called Xymasking. It is basically having strips of line across X or Y axis to block information </p>",
          "rawMarkdown": "They are time and frequency domain masking. On albumentation it's called Xymasking. It is basically having strips of line across X or Y axis to block information ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2814821,
      "postDate": "2024-05-15T15:11:10.840Z",
      "content": "<p>I’m interested, if no correlation is noticeable, then which notebook will you upload at the end, the one that scored the highest PB or CV?</p>",
      "rawMarkdown": "I’m interested, if no correlation is noticeable, then which notebook will you upload at the end, the one that scored the highest PB or CV?",
      "votes": 2
    },
    {
      "id": 2790001,
      "postDate": "2024-05-03T00:48:31.747Z",
      "content": "<p>What do you set alpha parameter to for FocalBCE loss </p>",
      "rawMarkdown": "What do you set alpha parameter to for FocalBCE loss ",
      "votes": 2,
      "replies": [
        {
          "id": 2791958,
          "postDate": "2024-05-04T00:45:42.300Z",
          "content": "<p>same with here: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/499713\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/499713</a></p>",
          "rawMarkdown": "same with here: https://www.kaggle.com/competitions/birdclef-2024/discussion/499713",
          "votes": 1
        }
      ]
    },
    {
      "id": 2778586,
      "postDate": "2024-04-27T08:43:52.867Z",
      "content": "<p>0.67 ensemble with a small model based on efficient net (0.66) , LB got 0.68</p>",
      "rawMarkdown": "0.67 ensemble with a small model based on efficient net (0.66) , LB got 0.68",
      "votes": 2
    },
    {
      "id": 2774215,
      "postDate": "2024-04-25T05:04:45.820Z",
      "content": "<p>Very interesting results, thanks for sharing.</p>\n<p>I wonder what is the CV change in the experiment with selecting either random crop or first 5 sec crop during training? <br>\nIn my experiments, random crop is always better :o</p>",
      "rawMarkdown": "Very interesting results, thanks for sharing.\n\nI wonder what is the CV change in the experiment with selecting either random crop or first 5 sec crop during training? \nIn my experiments, random crop is always better :o",
      "votes": 2,
      "replies": [
        {
          "id": 2774239,
          "postDate": "2024-04-25T05:16:05.020Z",
          "content": "<p>random crop is always better for me too in CV,  but LB is worse for me. you can see the interesting is that:<br>\nGoogle Model in 2024, if the first 5 seconds of the clip are selected, the metric is at 0.97325. If selected randomly, it is between 0.94-0.95. </p>",
          "rawMarkdown": "random crop is always better for me too in CV,  but LB is worse for me. you can see the interesting is that:\nGoogle Model in 2024, if the first 5 seconds of the clip are selected, the metric is at 0.97325. If selected randomly, it is between 0.94-0.95. \n",
          "votes": 4
        },
        {
          "id": 2774249,
          "postDate": "2024-04-25T05:23:35.133Z",
          "content": "<p>Random sampling for 5 seconds or only sampling the first 5 seconds might not be the best approach, as 5 seconds is too short and might crop to noise or silence. I am using the Google Bird Model to predict the scores of all segments and only randomly sample those with scores greater than a certain threshold.</p>",
          "rawMarkdown": "Random sampling for 5 seconds or only sampling the first 5 seconds might not be the best approach, as 5 seconds is too short and might crop to noise or silence. I am using the Google Bird Model to predict the scores of all segments and only randomly sample those with scores greater than a certain threshold.",
          "votes": 11,
          "replies": [
            {
              "id": 2774272,
              "postDate": "2024-04-25T05:50:24.417Z",
              "content": "<p>What is Google's Bird Model? Can you share some links please.</p>",
              "rawMarkdown": "What is Google's Bird Model? Can you share some links please.",
              "votes": 1
            },
            {
              "id": 2774292,
              "postDate": "2024-04-25T06:07:18.060Z",
              "content": "<p><a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier\" target=\"_blank\">https://www.kaggle.com/models/google/bird-vocalization-classifier</a></p>",
              "rawMarkdown": "https://www.kaggle.com/models/google/bird-vocalization-classifier",
              "votes": 4
            },
            {
              "id": 2778441,
              "postDate": "2024-04-27T07:03:57.003Z",
              "content": "<p>Thanks for the link to the model!</p>",
              "rawMarkdown": "Thanks for the link to the model!"
            },
            {
              "id": 2778492,
              "postDate": "2024-04-27T07:23:06.407Z",
              "content": "<p>By the way, do you know if this model exists in PyTorch (or at least a re-implementation)? Thanks!</p>",
              "rawMarkdown": "By the way, do you know if this model exists in PyTorch (or at least a re-implementation)? Thanks!",
              "votes": 3
            },
            {
              "id": 2781721,
              "postDate": "2024-04-29T00:54:02.617Z",
              "content": "<p>I didn't find PyTorch based (also want to fine-tune  hahaha),  I use tflite based to generate predictions</p>",
              "rawMarkdown": "I didn't find PyTorch based (also want to fine-tune  hahaha),  I use tflite based to generate predictions",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2774041,
      "postDate": "2024-04-25T03:26:34.873Z",
      "content": "<p>Thanks for sharing! How does your CV look for this? I see that you have shared a number of LB scores but I am curious how well your CV is correlating?</p>",
      "rawMarkdown": "Thanks for sharing! How does your CV look for this? I see that you have shared a number of LB scores but I am curious how well your CV is correlating?",
      "votes": 2,
      "replies": [
        {
          "id": 2775837,
          "postDate": "2024-04-25T20:40:27.193Z",
          "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>, <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> I am also very curious! I think setting up a stable and correlating CV is a big challenge in this competition since we do not have the true labels for every 5 seconds of audio data. That can get tricky for audio samples with a lot of secondary birds. </p>",
          "rawMarkdown": "@cody11null, @lihaoweicvch I am also very curious! I think setting up a stable and correlating CV is a big challenge in this competition since we do not have the true labels for every 5 seconds of audio data. That can get tricky for audio samples with a lot of secondary birds. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2773955,
      "postDate": "2024-04-25T01:51:26.910Z",
      "content": "<p>By first 5 sec do you mean only using first 5 sec of each ogg file for train data</p>",
      "rawMarkdown": "By first 5 sec do you mean only using first 5 sec of each ogg file for train data",
      "votes": 2,
      "replies": [
        {
          "id": 2773967,
          "postDate": "2024-04-25T02:06:25.887Z",
          "content": "<p>yes, using first 5 sec of each ogg file for train data</p>",
          "rawMarkdown": "yes, using first 5 sec of each ogg file for train data",
          "votes": 7,
          "replies": [
            {
              "id": 2778566,
              "postDate": "2024-04-27T08:32:46.503Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2776602,
      "postDate": "2024-04-26T08:46:49.247Z",
      "content": "<p>Thanks for sharing your insights. What about data? Do you use pretraining from previous years?</p>",
      "rawMarkdown": "Thanks for sharing your insights. What about data? Do you use pretraining from previous years?",
      "replies": [
        {
          "id": 2776607,
          "postDate": "2024-04-26T08:49:45.680Z",
          "content": "<p>No extra data now, I didn't use pretrained model from previous year,  just imagenet pretrained.<br>\nmy next step will be  pre-training and add more data from Xeno</p>",
          "rawMarkdown": "No extra data now, I didn't use pretrained model from previous year,  just imagenet pretrained.\nmy next step will be  pre-training and add more data from Xeno",
          "votes": 1,
          "replies": [
            {
              "id": 2776615,
              "postDate": "2024-04-26T08:53:18.357Z",
              "content": "<p>That's already a great performance, well done!</p>",
              "rawMarkdown": "That's already a great performance, well done!"
            }
          ]
        }
      ]
    },
    {
      "id": 2775311,
      "postDate": "2024-04-25T15:52:19.697Z",
      "content": "<p>Thank you for sharing! Did you use the knowledge distillation method proposed in the 4th place solution from 2023? </p>",
      "rawMarkdown": "Thank you for sharing! Did you use the knowledge distillation method proposed in the 4th place solution from 2023? ",
      "replies": [
        {
          "id": 2775992,
          "postDate": "2024-04-26T00:07:35.407Z",
          "content": "<p>I'm trying this and will update later</p>",
          "rawMarkdown": "I'm trying this and will update later",
          "votes": 1
        }
      ]
    },
    {
      "id": 2774977,
      "postDate": "2024-04-25T12:27:03.283Z",
      "content": "<p>yes, using first 5 sec of each ogg file for train data</p>",
      "rawMarkdown": "yes, using first 5 sec of each ogg file for train data"
    },
    {
      "id": 2774468,
      "postDate": "2024-04-25T07:52:44.893Z",
      "content": "<p>Thank you for sharing your results! What is your inference time per model and did you check what the ensemble results look like as well? Thanks.</p>",
      "rawMarkdown": "Thank you for sharing your results! What is your inference time per model and did you check what the ensemble results look like as well? Thanks.",
      "replies": [
        {
          "id": 2774629,
          "postDate": "2024-04-25T09:13:44.773Z",
          "content": "<p>single model takes about 1hour 20min with ONNX, so I did't try ensemble now</p>",
          "rawMarkdown": "single model takes about 1hour 20min with ONNX, so I did't try ensemble now",
          "votes": 4
        }
      ]
    },
    {
      "id": 2773980,
      "postDate": "2024-04-25T02:23:30.927Z",
      "content": "<p><code>LB: 0.66\nmodel2 + aug1 + only first 5sec (fold3 is 0.67)</code></p>\n<p>Why did you choose a specific fold in 5-folds? Is that just because its validation score is the best?</p>",
      "rawMarkdown": "`LB: 0.66\nmodel2 + aug1 + only first 5sec (fold3 is 0.67)`\n\nWhy did you choose a specific fold in 5-folds? Is that just because its validation score is the best?",
      "replies": [
        {
          "id": 2773997,
          "postDate": "2024-04-25T02:46:24.017Z",
          "content": "<p>I choose fold2 for compare<br>\nLB score on my side:<br>\nfold1:  0.64<br>\nfold2:  0.66<br>\nfold3:  0.67<br>\nfold4:  0.64<br>\nfold5:  0.64</p>",
          "rawMarkdown": "I choose fold2 for compare\nLB score on my side:\nfold1:  0.64\nfold2:  0.66\nfold3:  0.67\nfold4:  0.64\nfold5:  0.64",
          "votes": 1,
          "replies": [
            {
              "id": 2774141,
              "postDate": "2024-04-25T04:21:55.677Z",
              "content": "<p>Hi, thanks for sharing!<br>\nIs this result correlating to valid score ?</p>",
              "rawMarkdown": "Hi, thanks for sharing!\nIs this result correlating to valid score ?"
            }
          ]
        }
      ]
    },
    {
      "id": 2781788,
      "postDate": "2024-04-29T02:31:10.113Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2783966,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-30T03:24:08.210000",
      "content": "<p>Update 04-30,  Here is my latest experiments (fig): <br>\nI tried pretrain,  add nocall,  train with extend class,  multi audio clip with score …<br>\nStill hard for me to find the correlation!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F2763710548b0f26f78580c547acaffaa%2F6951714446627_.pic.jpg?generation=1714447369360418&amp;alt=media\"></p>",
      "votes": 9,
      "replies": [
        {
          "id": 2783976,
          "author_name": "LLLEEEOOOH",
          "author_url": "",
          "post_date": "2024-04-30T03:34:21.340000",
          "content": "<p>Thank you for sharing! May I ask what does cfg0/1/2/3 means?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2783981,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-30T03:47:37.687000",
              "content": "<p>cfg0 is my baseline with label_smooth=0, second coef=1.0  and some augs<br>\ncfg1～cfg3 is diffrent second coef mainly (poor so I didn't submit all)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2783999,
              "author_name": "LLLEEEOOOH",
              "author_url": "",
              "post_date": "2024-04-30T04:07:55.583000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2785750,
          "author_name": "Zhongkai Shangguan",
          "author_url": "",
          "post_date": "2024-05-01T02:43:53.180000",
          "content": "<p>wow this is super impressive! may I ask which backbone are you using?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2785758,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-05-01T02:55:12.480000",
              "content": "<p>I’m using eca_nfnet_l0</p>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 2787737,
          "author_name": "Sinan Calisir",
          "author_url": "",
          "post_date": "2024-05-01T23:18:49.907000",
          "content": "<p>Interesting that the pretraining did not improve the scores?! Did you able to figure it out? In my case, with 5 epochs pretraining with the previous comp data, it gives some boost.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2791956,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-05-04T00:44:34.233000",
              "content": "<p>I'm debugging this, strange for me it gives no boost. 🤣</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2780851,
      "author_name": "Salman Ahmed",
      "author_url": "",
      "post_date": "2024-04-28T13:03:20.260000",
      "content": "<p>Well for the same experiment settings, I am unable to reproduce these results. Following 2023 4th place solution.<br>\nFirst 5 seconds training. </p>\n<p>Apparently, the only changes I can think of might be the followings: <br>\nEarly Stopping (Around 40 Epochs)? (As the validation set is not reliable, so small reduction in ROC or CMAP is not a good way to early stop.)<br>\nTime Mask / Freq Mask?<br>\neca_nfnet_l0 <br>\nOnnx inference<br>\nNormalizing the Mel Spec or Normalization of Wave?<br>\nMixed Precision?<br>\nGoogle Bird Model Predictions for KD (Version 2)</p>\n<p>Not sure, if I missed something. Remaining setting is same. (I am only able to get the score around 0.62 max)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2781086,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-04-28T14:53:55.437000",
          "content": "<p>Did you add the thresholding mentioned <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/497539#2774249\" target=\"_blank\">here</a>? I think this might be a key component.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2781116,
              "author_name": "Shiro",
              "author_url": "",
              "post_date": "2024-04-28T15:17:23.497000",
              "content": "<p>good point i miss this details. i'm also struggling to reproduce results.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2781712,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-29T00:44:37.827000",
          "content": "<ol>\n<li>Time Mask / Freq Mask? --- Time and Freq</li>\n<li>Normalizing the Mel Spec or Normalization of Wave? -- On Mel spec, norm to 0~255 then normalize by imagenet RGB mean std.</li>\n<li>Mixed Precision? -- yes</li>\n</ol>\n<p>Based on the information you provided, I suggest you to check:</p>\n<p>Using FocalBCE<br>\nChecking sencond_coef, label_smooth</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2781727,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-29T01:03:02.267000",
              "content": "<p><a href=\"https://www.kaggle.com/salmanahmedtamu\" target=\"_blank\">@salmanahmedtamu</a>   and no KD (this slightly decreased my cv), sharing more detailed information about your CV might be beneficial for analysis.  my CV range from  0.972 -- 0.974 </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781823,
              "author_name": "Salman Ahmed",
              "author_url": "",
              "post_date": "2024-04-29T02:56:48.237000",
              "content": "<p>Yeah, I meant did you use Time and Freq Augs?</p>\n<p>My CV ranges from 0.98 to 0.99</p>\n<p>I used label smooth to 0.05, and secondary_coefs to 1.0</p>\n<p>0.66 or 0.67 is without KD? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781825,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-29T02:58:23.810000",
              "content": "<p>yes I used.  without KD and without pretrain,use imagenet pretrained,  0.98 to 0.99 is very high, I guess your validation don't use secondary_coef 1.0 ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781829,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-29T03:02:44.960000",
              "content": "<p>and my fold2 log:<br>\nsteps: 395,epo: 1, best_score: 0.923567 train_loss: 0.029568 , eval_loss: 0.023104946649113257 , roc_auc: 0.9235667694966543 <br>\nsteps: 790,epo: 2, best_score: 0.958245 train_loss: 0.018857 , eval_loss: 0.017212639977042744 , roc_auc: 0.9582446547051086 <br>\nsteps: 1185,epo: 3, best_score: 0.966371 train_loss: 0.016296 , eval_loss: 0.015339857312040283 , roc_auc: 0.9663713526370147 <br>\nsteps: 1580,epo: 4, best_score: 0.969536 train_loss: 0.014625 , eval_loss: 0.014284247196847935 , roc_auc: 0.9695358129988351 <br>\nsteps: 1975,epo: 5, best_score: 0.971608 train_loss: 0.013271 , eval_loss: 0.013741090936078266 , roc_auc: 0.9716079793856934 <br>\nsteps: 2370,epo: 6, best_score: 0.974090 train_loss: 0.010935 , eval_loss: 0.013370304558934136 , roc_auc: 0.9740900915645582 <br>\nsteps: 2765,epo: 7, best_score: 0.974090 train_loss: 0.009982 , eval_loss: 0.013458995466035876 , roc_auc: 0.9736567230864034 <br>\nsteps: 3160,epo: 8, best_score: 0.974512 train_loss: 0.009245 , eval_loss: 0.013592320273618227 , roc_auc: 0.9745117067793344 <br>\nsteps: 3555,epo: 9, best_score: 0.974512 train_loss: 0.008837 , eval_loss: 0.013835379566586056 , roc_auc: 0.9740697744987592 <br>\nsteps: 3950,epo: 10, best_score: 0.974512 train_loss: 0.008407 , eval_loss: 0.014064931419569177 , roc_auc: 0.9725833905498832 <br>\nsteps: 4345,epo: 11, best_score: 0.974512 train_loss: 0.007450 , eval_loss: 0.015370996599388006 , roc_auc: 0.9704292221454025 <br>\nsteps: 4740,epo: 12, best_score: 0.974512 train_loss: 0.006486 , eval_loss: 0.016247244984177605 , roc_auc: 0.9698166650036316 <br>\nsteps: 5135,epo: 13, best_score: 0.974512 train_loss: 0.006573 , eval_loss: 0.016679136829402346 , roc_auc: 0.9687894194276919 <br>\n2024-04-15_18:13:52: early stopped!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781830,
              "author_name": "Salman Ahmed",
              "author_url": "",
              "post_date": "2024-04-29T03:03:17.320000",
              "content": "<p>Yeah, I don't have secondary_coefs in validation.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781834,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-29T03:08:21.923000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781838,
              "author_name": "Salman Ahmed",
              "author_url": "",
              "post_date": "2024-04-29T03:15:28.583000",
              "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Do you use sampling, or weights for the loss?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2781846,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-29T03:18:05.493000",
              "content": "<p>currently no, but I'll add this later, sampling and weighting by rating</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2782114,
              "author_name": "Salman Ahmed",
              "author_url": "",
              "post_date": "2024-04-29T06:17:17.897000",
              "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> I added secondary labels in validation and used focalbce. <br>\nNow my CV is 0.9575</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2783700,
              "author_name": "Shiro",
              "author_url": "",
              "post_date": "2024-04-29T22:13:29.320000",
              "content": "<p>what is secondary_coef? are you using the secondary label as birds to detect ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2783756,
              "author_name": "Salman Ahmed",
              "author_url": "",
              "post_date": "2024-04-29T23:52:02.460000",
              "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> It means, target value for secondary labels, if its going to be 1 or 0.8 or whatever you want to set. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2786180,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-05-01T07:48:40.137000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2773935,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-25T01:19:09.330000",
      "content": "<p>Now my focus work is to explore the influence of the data quality, follows are doing:</p>\n<ul>\n<li>use trained model select out the high quality segs by predicted score, use Google Kaggle Bird Model maybe be a good choice?</li>\n</ul>\n<p>According to the training data prediction of Google Model in 2024, if the first 5 seconds of the clip are selected, the metric is at 0.97325. If selected randomly, it is between 0.94-0.95. However, if the segment with the highest category score is chosen, it goes up to 0.983.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2781790,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-29T02:31:59.587000",
      "content": "<p>Here is my latest experiments (fig): <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F1dbd036d6cab4700009843ef7503dbea%2F6911714357731_.pic.jpg?generation=1714357891979872&amp;alt=media\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2776589,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-26T08:40:49.267000",
      "content": "<p>Has anyone achieved good results with ViT? （figure below is  my results using different backbone）In addition, in this competition, we need to balance cost-effectiveness. It is possible that an ensemble of smaller models may win.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F35c834f2a37f6f3fd44c07b1e3c1d37f%2F6891714120735_.pic.jpg?generation=1714120847514341&amp;alt=media\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 2776790,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-26T10:37:51.300000",
          "content": "<p>efficientvit_b2  has export Bug, the LB score is not correct, I'll update later</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2777069,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-04-26T13:55:37.833000",
          "content": "<p>I started with <code>efficientvit</code>, but moved to <code>efficientnet</code> as the models can be converted to onnx without modifying the padding layers. Also, in my early experiments, CV/LB was stronger with <code>efficientnet</code>.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2778162,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-27T02:07:00.153000",
              "content": "<p>After fixed the bug, my efficienvit got 0.65 which is also slightly lower than efficientnet.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2780018,
          "author_name": "deepshark",
          "author_url": "",
          "post_date": "2024-04-28T01:41:08.017000",
          "content": "<p>what is the meaning of LB？</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2780049,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-28T02:32:34.780000",
              "content": "<p>Leaderboard,  Public  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2776080,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-26T01:20:04.860000",
      "content": "<p>I have a question: in the online test set, would a 5-second segment include multiple birds? If there are multiple birds, what coefficient would be used? Is it 1?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2777049,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2024-04-26T13:37:32.190000",
          "content": "<p>Yes, there can be multiple birds, yes all birds that are audible in the 5-second segment have the ground truth 1</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2778164,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-27T02:07:46.780000",
              "content": "<p>Thank you for sharing </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2784770,
              "author_name": "Will Rice",
              "author_url": "",
              "post_date": "2024-04-30T13:37:45.667000",
              "content": "<p>Does that mean a multiclass approach is a dead end?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2787866,
              "author_name": "Shadi Akiki",
              "author_url": "",
              "post_date": "2024-05-02T01:14:39.230000",
              "content": "<blockquote>\n  <p>all birds that are audible in the 5-second segment have the ground truth 1</p>\n</blockquote>\n<p>What makes you think so?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2787964,
              "author_name": "Will Rice",
              "author_url": "",
              "post_date": "2024-05-02T03:31:23.673000",
              "content": "<p>Because that describes a multilabel problem.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2789352,
              "author_name": "Will Rice",
              "author_url": "",
              "post_date": "2024-05-02T16:41:56.643000",
              "content": "<p>I am also struggling to reach my same LB performance after using FocalBCE instead of CrossEntropy though.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2774228,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-25T05:11:20.933000",
      "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>  <a href=\"https://www.kaggle.com/lmhongkhnh\" target=\"_blank\">@lmhongkhnh</a>  I didn't find correlation, even though CV got  0.985 (with KD), LB may get  0.64<br>\nfold1:  CV: 0.974080, LB:  0.64<br>\nfold2:  CV: 0.974512, LB:  0.66<br>\nfold3:  CV: 0.971235, LB:  0.67<br>\nfold4:  CV: 0.973694, LB:  0.64<br>\nfold5: CV: 0.970303, LB:  0.64</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2775234,
          "author_name": "Hugo de Heer",
          "author_url": "",
          "post_date": "2024-04-25T14:56:46.343000",
          "content": "<p>How did you setup your CV? Only on primary labels? And did you evaluate on samples &gt;= than a specific quality rating? Or on all data. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2775998,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-26T00:11:41.990000",
              "content": "<p>I used secondary_coef = 1.0 and evaluate on all without specific quality rating.<br>\nHowever, I don't think this is the most reasonable approach, we need more experiments, especially when there are multiple tags. Taking only 5 seconds is obviously not necessarily enough to capture multiple birds calling.</p>\n<pre><code> s  secondary_labels:\n       s !=  and s  CFG():\n             target] = self\n</code></pre>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2799893,
      "author_name": "ShawnTung",
      "author_url": "",
      "post_date": "2024-05-08T02:53:01.450000",
      "content": "<p>Nice work! May I ask how you ensemble? When I try to use concurrent.futures.ThreadPoolExecutor to perform inference on multiple models simultaneously, the execution runs smoothly but errors occur during submission.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2802272,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-05-09T00:41:45.063000",
          "content": "<p>try batch_size=48  with onnx session option num_threads=4</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2803250,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-05-09T11:58:39.563000",
              "content": "<p>My onnx is slower than jit with batch_size=4, with batch_size=48 it could be faster?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2803406,
              "author_name": "ShawnTung",
              "author_url": "",
              "post_date": "2024-05-09T13:42:13.290000",
              "content": "<p>I don't know why all my submissions using num_threads=4 have failed….</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2795816,
      "author_name": "Arun",
      "author_url": "",
      "post_date": "2024-05-06T02:19:32.993000",
      "content": "<p>Have you tried efficientnet models ? How well do they perform for you ? I tried comparing few models like efficientnet, mobile net and efficientnet v2, there isn't much significant change when I performed them on the same pipeline,  what's your results on this ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2796122,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-05-06T06:04:21.510000",
          "content": "<p>I tried efficientnet b0 to b2 and eca_nfnet_l0,   different folds perform differently, there's no absolute best across different  backbone, I usually fluctuate between 0.64--0.68. Your situation seems nice, pretty stable, could you share more about your data pipeline if you don't mind?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2796222,
              "author_name": "Arun",
              "author_url": "",
              "post_date": "2024-05-06T06:43:17.260000",
              "content": "<p>tbh, my situation isn't stable either, I cannot recreate my best scoring model myself, a lot of variables and randomness is involved, and i am not using the folds strategy, the cv isn't so reliable anyway so I am using most of the data chunk with a little bit of cleaning  in the training process</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2796737,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-05-06T11:39:16.543000",
              "content": "<p>Thank you for sharing， it is the same case on my side， LB is very sensitive </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2803251,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2024-05-09T11:59:38.333000",
              "content": "<p>It seems that same CV for different models do not correlate with LB score?</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2782100,
      "author_name": "zz1zzi",
      "author_url": "",
      "post_date": "2024-04-29T06:00:27.380000",
      "content": "<p>Thank you for sharing your solution. Could you please clarify how you transform the waveform into a MelSpectrogram, whether it's done online or offline? Additionally, the 4th place solution in 2023 transforms the waveform into a MelSpectrogram during the forward propagation, for instance, e.g., </p>\n<pre><code>     ():\n        x = []\n         CFG.use_pcen:\n            x = self.pcen(x).unsqueeze()\n            self.pcen.reset()\n        :\n            x = self.logmelspec_extractor(x)[:, ] \n\n        \n        \n        \n        \n        \n        \n</code></pre>\n<p>How can I incorporate data transformation using Albumentations in the model's forward pass? Thanks a lot.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2782423,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-29T09:52:29.747000",
          "content": "<p>you can try librosa.melspec in dataset</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2781897,
      "author_name": "khushiarora09",
      "author_url": "",
      "post_date": "2024-04-29T04:06:15.403000",
      "content": "<p>nicee work </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2781863,
      "author_name": "Salman Ahmed",
      "author_url": "",
      "post_date": "2024-04-29T03:25:26.037000",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Imagenet mean and std are something like this. <br>\nnormalize = transforms.Normalize(mean=[0.485, 0.456, 0.406],<br>\n                                 std=[0.229, 0.224, 0.225])</p>\n<p>I have 2 assumptions here.</p>\n<ol>\n<li><p>you multiplied these mean and std by 255.0 and then normalized. <br>\nWhich means you input to the model might not be in range of 0 - 1</p></li>\n<li><p>You normalize to 0-1 and then applied this imagenet norm.</p></li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 2781870,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-29T03:29:25.633000",
          "content": "<p>yes you are right,  like this, input image is normalize to 0~255 first</p>\n<pre><code>def normalize(img, , , max_pixel_value=):\n     = .(, dtype=.float32)\n     *= max_pixel_value\n\n     = .(, dtype=.float32)\n     *= max_pixel_value\n\n    denominator = .reciprocal(, dtype=.float32)\n\n    img = img.astype(.float32)\n    img -= \n    img *= denominator\n     img\n</code></pre>",
          "votes": 1,
          "replies": [
            {
              "id": 2787565,
              "author_name": "Aurelio",
              "author_url": "",
              "post_date": "2024-05-01T20:06:54.650000",
              "content": "<p>I'm a bit confused on what you do here, can't you just use np.interp to get the values of the image directly to [0,1] and then do the imagenet normalization? Also one curiosity, do you use wighted sampling for your dataloader?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2791959,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-05-04T00:47:27.397000",
              "content": "<p>weighted sampling  low down my score, see my latest experiments </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2781767,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-29T02:15:21.593000",
      "content": "<p>Here is a question: If we have trained with more data segments and have observed significant improvement locally, would the recall rate usually increase when applied to another dataset?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2781662,
      "author_name": "thacrobatheskis",
      "author_url": "",
      "post_date": "2024-04-28T23:11:20.323000",
      "content": "<p>Thank you for posting your findings!</p>\n<blockquote>\n  <p>5fold, Focal BCE, EMA</p>\n</blockquote>\n<p>What is EMA?</p>\n<blockquote>\n  <p>LB: 0.63</p>\n</blockquote>\n<p>How do you compute the LB score for a model that had 5-fold CV? Do you submit all 5 models then average their LBs, or create 1 submission from an ensemble of all 5 models, or just submit 1 of the 5 that performed the best?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2781706,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-29T00:39:57.707000",
          "content": "<p>EMA： Exponential Moving Average</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2781709,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-29T00:41:24.633000",
              "content": "<p>I submit 5 folds with model2 (LB  0.64~0.67),    model1  I only  submit fold2</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2781614,
      "author_name": "Manav2805",
      "author_url": "",
      "post_date": "2024-04-28T21:25:49.663000",
      "content": "<p>Should we try to use the centre 5 seconds of the audio clipping rather than the first 5 seconds. It is just an intuitive thought. Will this work or is there something fundamentally wrong with it ??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2781648,
          "author_name": "snehal",
          "author_url": "",
          "post_date": "2024-04-28T23:00:28.100000",
          "content": "<p>We want to avoid using 5 second clips that have no birdcall in our train data but are labeled as having some call. </p>\n<p>If a recording in the train set is 15 seconds and the label is pigeon it means somewhere in the 15 seconds there was a pigeon call. If we think about how the data was created it seems likely that this 15 second audio clip was taken from a longer recording at some location. When clipping the audio it makes sense that the start and end of the clip would be made near a pigeon call (why would you clip a recording to start 30 seconds before the first birdcall or end 30 seconds after any birdcall). Therefore first 5 seconds and last 5 seconds of any train data seem to be logical choices for less noisy data (but still possible that they have no birdcall).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2781711,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-29T00:44:04.953000",
          "content": "<p>you can try center 5 seconds and maybe share to us, but I think  finding the audio clip which really contain bird calls can be more important.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2780836,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2024-04-28T12:56:58.590000",
      "content": "<p>Thanks for sharing again! Do you train only with the comp data and for how many epochs are you training for?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2781713,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-29T00:48:35.010000",
          "content": "<p>Yes now I only use 2024 comp data,   train with EarlyStopping, usually  7-10 epoch with get the best </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2780830,
      "author_name": "Salman Ahmed",
      "author_url": "",
      "post_date": "2024-04-28T12:55:19.640000",
      "content": "<p>How many epochs have you been training for?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2781716,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-29T00:50:58.280000",
          "content": "<p>I train max 50 epochs, but train with EarlyStopping, usually 7-10 epoch with get the best, and I always submit the best</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2796311,
              "author_name": "thacrobatheskis",
              "author_url": "",
              "post_date": "2024-05-06T07:18:18.827000",
              "content": "<p>How do you determine the best? (I’m noticing that because local validation is unreliable, I don’t know when to stop training.)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2780069,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-28T03:02:35.693000",
      "content": "<p>Using multiple CVs to examine the correlation with LB may make it easier to discover the correlation ？</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0cb83b7c376ff5c13ae0b637dbe6a40a%2F6901714272397_.pic.jpg?generation=1714273338317490&amp;alt=media\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2781074,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2024-04-28T14:44:31.673000",
          "content": "<p>question  about your training, about the model , are you applying also the Kownledge Distillation and augmentation mentionned in the 4th solution? or are you just using the same backbone?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2781718,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-29T00:51:56.673000",
              "content": "<p>I  only use the same arch and backbone, no KD now</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2777565,
      "author_name": "HB",
      "author_url": "",
      "post_date": "2024-04-26T17:54:45.460000",
      "content": "<p>Could you tell me the score change when you used your notebook Bird Sound Denoise by Deep Model for train data and test data? And I wonder if you have tested 7 seconds, 10 seconds, etc. instead of the first 5 seconds.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2778167,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-27T02:09:28.177000",
          "content": "<p>I didn’t use denoise now because that model hurt the bird sound. 7sec 10sec you mean train or infer ? </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2778567,
              "author_name": "HB",
              "author_url": "",
              "post_date": "2024-04-27T08:33:01.063000",
              "content": "<p>There will be damage to the sound, but I'm quite curious about the result. I was talking about the case of training. I didn't see the increase in inference well, but I wondered what would happen to inference as well.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2778575,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-27T08:38:12.120000",
              "content": "<p>I tried 10s train in the early stage, then infer with 5s LB 0.64,  infer with 10s repeating to 5s LB 0.6</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2778679,
              "author_name": "HB",
              "author_url": "",
              "post_date": "2024-04-27T09:22:58.593000",
              "content": "<p>Thank you for your kind reply!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2775698,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2024-04-25T18:42:03.163000",
      "content": "<p>interesting, thanks for sharing;</p>\n<p>what kind of augmentation are you using regarding \"use Freq and Time aug, \". i try to use somes but haven't see big diff</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2775723,
          "author_name": "LLLEEEOOOH",
          "author_url": "",
          "post_date": "2024-04-25T18:52:21.220000",
          "content": "<p>They are time and frequency domain masking. On albumentation it's called Xymasking. It is basically having strips of line across X or Y axis to block information </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2814821,
      "author_name": "Karnakbayev Artur",
      "author_url": "",
      "post_date": "2024-05-15T15:11:10.840000",
      "content": "<p>I’m interested, if no correlation is noticeable, then which notebook will you upload at the end, the one that scored the highest PB or CV?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2790001,
      "author_name": "snehal",
      "author_url": "",
      "post_date": "2024-05-03T00:48:31.747000",
      "content": "<p>What do you set alpha parameter to for FocalBCE loss </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2791958,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-05-04T00:45:42.300000",
          "content": "<p>same with here: <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/499713\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/499713</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2778586,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-27T08:43:52.867000",
      "content": "<p>0.67 ensemble with a small model based on efficient net (0.66) , LB got 0.68</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2774215,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-25T05:04:45.820000",
      "content": "<p>Very interesting results, thanks for sharing.</p>\n<p>I wonder what is the CV change in the experiment with selecting either random crop or first 5 sec crop during training? <br>\nIn my experiments, random crop is always better :o</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2774239,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-25T05:16:05.020000",
          "content": "<p>random crop is always better for me too in CV,  but LB is worse for me. you can see the interesting is that:<br>\nGoogle Model in 2024, if the first 5 seconds of the clip are selected, the metric is at 0.97325. If selected randomly, it is between 0.94-0.95. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2774249,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-25T05:23:35.133000",
          "content": "<p>Random sampling for 5 seconds or only sampling the first 5 seconds might not be the best approach, as 5 seconds is too short and might crop to noise or silence. I am using the Google Bird Model to predict the scores of all segments and only randomly sample those with scores greater than a certain threshold.</p>",
          "votes": 11,
          "replies": [
            {
              "id": 2774272,
              "author_name": "Phaedrus",
              "author_url": "",
              "post_date": "2024-04-25T05:50:24.417000",
              "content": "<p>What is Google's Bird Model? Can you share some links please.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2774292,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-25T06:07:18.060000",
              "content": "<p><a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier\" target=\"_blank\">https://www.kaggle.com/models/google/bird-vocalization-classifier</a></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2778441,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2024-04-27T07:03:57.003000",
              "content": "<p>Thanks for the link to the model!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2778492,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2024-04-27T07:23:06.407000",
              "content": "<p>By the way, do you know if this model exists in PyTorch (or at least a re-implementation)? Thanks!</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2781721,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-04-29T00:54:02.617000",
              "content": "<p>I didn't find PyTorch based (also want to fine-tune  hahaha),  I use tflite based to generate predictions</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2774041,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2024-04-25T03:26:34.873000",
      "content": "<p>Thanks for sharing! How does your CV look for this? I see that you have shared a number of LB scores but I am curious how well your CV is correlating?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2775837,
          "author_name": "Hugo de Heer",
          "author_url": "",
          "post_date": "2024-04-25T20:40:27.193000",
          "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>, <a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> I am also very curious! I think setting up a stable and correlating CV is a big challenge in this competition since we do not have the true labels for every 5 seconds of audio data. That can get tricky for audio samples with a lot of secondary birds. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2773955,
      "author_name": "snehal",
      "author_url": "",
      "post_date": "2024-04-25T01:51:26.910000",
      "content": "<p>By first 5 sec do you mean only using first 5 sec of each ogg file for train data</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2773967,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-25T02:06:25.887000",
          "content": "<p>yes, using first 5 sec of each ogg file for train data</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2778566,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-27T08:32:46.503000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2776602,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2024-04-26T08:46:49.247000",
      "content": "<p>Thanks for sharing your insights. What about data? Do you use pretraining from previous years?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2776607,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-26T08:49:45.680000",
          "content": "<p>No extra data now, I didn't use pretrained model from previous year,  just imagenet pretrained.<br>\nmy next step will be  pre-training and add more data from Xeno</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2776615,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2024-04-26T08:53:18.357000",
              "content": "<p>That's already a great performance, well done!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2775311,
      "author_name": "LLLEEEOOOH",
      "author_url": "",
      "post_date": "2024-04-25T15:52:19.697000",
      "content": "<p>Thank you for sharing! Did you use the knowledge distillation method proposed in the 4th place solution from 2023? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2775992,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-04-26T00:07:35.407000",
          "content": "<p>I'm trying this and will update later</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2774977,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-25T12:27:03.283000",
      "content": "<p>yes, using first 5 sec of each ogg file for train data</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2774468,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T07:52:44.893000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2774629,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-25T09:13:44.773000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2773980,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T02:23:30.927000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2773997,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-25T02:46:24.017000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2774141,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T04:21:55.677000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2781788,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-29T02:31:10.113000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2773933": "Basic Setting\n\n5fold,  Focal BCE, EMA\n\nmodel1:  2023 rank1 model\nmodel2:  2023 rank4 model\naug1:    norm to 0~255 then norm by imagenet RGB std,  then use Hflip, CoarseDropOut, Mixup\naug2:   norm to 0~1,  one channel, use  Freq and Time aug,  Mixup\n\nLB:  0.63\nmodel1 + aug2 + only first 5sec + with additional time domain noise aug \nLB:  0.64\nmodel1 + aug2 + only first 5sec\nLB:  0.62\nmodel1 + aug1 + only first 5sec  \n\n!! noise aug maybe harmful when data quality is low\n\nLB:  0.66  \nmodel2 + aug1 + only first 5sec   (fold3 is  0.67)\nLB:  0.6\nmodel2 + aug1 + random crop 5sec when train\n\n!!  random crop may crop just noise, and overfit the noise\nand I found that the data quality is very low, e.g  brfowl1  mix other bird sound\n\nFirst 5sec is good but not the best\n\nmodel2 all fold:\n\nfold1: CV: 0.974080, LB: 0.64\nfold2: CV: 0.974512, LB: 0.66\nfold3: CV: 0.971235, LB: 0.67\nfold4: CV: 0.973694, LB: 0.64\nfold5: CV: 0.970303, LB: 0.64",
    "2783966": "Update 04-30,  Here is my latest experiments (fig): \nI tried pretrain,  add nocall,  train with extend class,  multi audio clip with score ...\nStill hard for me to find the correlation!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F2763710548b0f26f78580c547acaffaa%2F6951714446627_.pic.jpg?generation=1714447369360418&alt=media)",
    "2780851": "Well for the same experiment settings, I am unable to reproduce these results. Following 2023 4th place solution.\nFirst 5 seconds training. \n\nApparently, the only changes I can think of might be the followings: \nEarly Stopping (Around 40 Epochs)? (As the validation set is not reliable, so small reduction in ROC or CMAP is not a good way to early stop.)\nTime Mask / Freq Mask?\neca_nfnet_l0 \nOnnx inference\nNormalizing the Mel Spec or Normalization of Wave?\nMixed Precision?\nGoogle Bird Model Predictions for KD (Version 2)\n\nNot sure, if I missed something. Remaining setting is same. (I am only able to get the score around 0.62 max)",
    "2773935": "Now my focus work is to explore the influence of the data quality, follows are doing:\n- use trained model select out the high quality segs by predicted score, use Google Kaggle Bird Model maybe be a good choice?\n\nAccording to the training data prediction of Google Model in 2024, if the first 5 seconds of the clip are selected, the metric is at 0.97325. If selected randomly, it is between 0.94-0.95. However, if the segment with the highest category score is chosen, it goes up to 0.983.\n\n",
    "2781790": "Here is my latest experiments (fig): \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F1dbd036d6cab4700009843ef7503dbea%2F6911714357731_.pic.jpg?generation=1714357891979872&alt=media)",
    "2776589": "Has anyone achieved good results with ViT? （figure below is  my results using different backbone）In addition, in this competition, we need to balance cost-effectiveness. It is possible that an ensemble of smaller models may win.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F35c834f2a37f6f3fd44c07b1e3c1d37f%2F6891714120735_.pic.jpg?generation=1714120847514341&alt=media)",
    "2776080": "I have a question: in the online test set, would a 5-second segment include multiple birds? If there are multiple birds, what coefficient would be used? Is it 1?",
    "2774228": "@cody11null  @lmhongkhnh  I didn't find correlation, even though CV got  0.985 (with KD), LB may get  0.64\nfold1:  CV: 0.974080, LB:  0.64\nfold2:  CV: 0.974512, LB:  0.66\nfold3:  CV: 0.971235, LB:  0.67\nfold4:  CV: 0.973694, LB:  0.64\nfold5: CV: 0.970303, LB:  0.64\n",
    "2799893": "Nice work! May I ask how you ensemble? When I try to use concurrent.futures.ThreadPoolExecutor to perform inference on multiple models simultaneously, the execution runs smoothly but errors occur during submission.",
    "2795816": "Have you tried efficientnet models ? How well do they perform for you ? I tried comparing few models like efficientnet, mobile net and efficientnet v2, there isn't much significant change when I performed them on the same pipeline,  what's your results on this ? ",
    "2782100": "Thank you for sharing your solution. Could you please clarify how you transform the waveform into a MelSpectrogram, whether it's done online or offline? Additionally, the 4th place solution in 2023 transforms the waveform into a MelSpectrogram during the forward propagation, for instance, e.g., \n```python\n    def forward(self, input):\n        x = input['wave']\n        if CFG.use_pcen:\n            x = self.pcen(x).unsqueeze(1)\n            self.pcen.reset()\n        else:\n            x = self.logmelspec_extractor(x)[:, None] # (32, 1, 128, 313)\n        \n        # TODO\n        # (32, 1, 128, 313) -> (32, 3, 128, 313)\n        # norm by imagenet RGB std\n        # then use Hflip\n        # CoarseDropOut\n        # Mixup\n```\n\nHow can I incorporate data transformation using Albumentations in the model's forward pass? Thanks a lot.",
    "2781897": "nicee work ",
    "2781863": "@lihaoweicvch Imagenet mean and std are something like this. \nnormalize = transforms.Normalize(mean=[0.485, 0.456, 0.406],\n                                 std=[0.229, 0.224, 0.225])\n\nI have 2 assumptions here.\n1. you multiplied these mean and std by 255.0 and then normalized. \nWhich means you input to the model might not be in range of 0 - 1\n\n2. You normalize to 0-1 and then applied this imagenet norm.",
    "2781767": "Here is a question: If we have trained with more data segments and have observed significant improvement locally, would the recall rate usually increase when applied to another dataset?",
    "2781662": "Thank you for posting your findings!\n\n> 5fold, Focal BCE, EMA\n\nWhat is EMA?\n\n> LB: 0.63\n\nHow do you compute the LB score for a model that had 5-fold CV? Do you submit all 5 models then average their LBs, or create 1 submission from an ensemble of all 5 models, or just submit 1 of the 5 that performed the best?",
    "2781614": "Should we try to use the centre 5 seconds of the audio clipping rather than the first 5 seconds. It is just an intuitive thought. Will this work or is there something fundamentally wrong with it ??",
    "2780836": "Thanks for sharing again! Do you train only with the comp data and for how many epochs are you training for?",
    "2780830": "How many epochs have you been training for?\n",
    "2780069": "Using multiple CVs to examine the correlation with LB may make it easier to discover the correlation ？\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0cb83b7c376ff5c13ae0b637dbe6a40a%2F6901714272397_.pic.jpg?generation=1714273338317490&alt=media)",
    "2777565": "Could you tell me the score change when you used your notebook Bird Sound Denoise by Deep Model for train data and test data? And I wonder if you have tested 7 seconds, 10 seconds, etc. instead of the first 5 seconds.",
    "2775698": "interesting, thanks for sharing;\n\nwhat kind of augmentation are you using regarding \"use Freq and Time aug, \". i try to use somes but haven't see big diff",
    "2814821": "I’m interested, if no correlation is noticeable, then which notebook will you upload at the end, the one that scored the highest PB or CV?",
    "2790001": "What do you set alpha parameter to for FocalBCE loss ",
    "2778586": "0.67 ensemble with a small model based on efficient net (0.66) , LB got 0.68",
    "2774215": "Very interesting results, thanks for sharing.\n\nI wonder what is the CV change in the experiment with selecting either random crop or first 5 sec crop during training? \nIn my experiments, random crop is always better :o",
    "2774041": "Thanks for sharing! How does your CV look for this? I see that you have shared a number of LB scores but I am curious how well your CV is correlating?",
    "2773955": "By first 5 sec do you mean only using first 5 sec of each ogg file for train data",
    "2776602": "Thanks for sharing your insights. What about data? Do you use pretraining from previous years?",
    "2775311": "Thank you for sharing! Did you use the knowledge distillation method proposed in the 4th place solution from 2023? ",
    "2774977": "yes, using first 5 sec of each ogg file for train data",
    "2774468": "Thank you for sharing your results! What is your inference time per model and did you check what the ensemble results look like as well? Thanks.",
    "2773980": "`LB: 0.66\nmodel2 + aug1 + only first 5sec (fold3 is 0.67)`\n\nWhy did you choose a specific fold in 5-folds? Is that just because its validation score is the best?",
    "2781788": ""
  }
}