{
  "id": 511559,
  "title": "Public 20th | Private 15th solution ",
  "url": "/competitions/birdclef-2024/writeups/innervoice-public-20th-private-15th-solution",
  "author_name": "",
  "post_date": "2024-06-11T08:57:41.653Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to competition host for hosting this year BirdClef a bit different  as compared to previous competitions. Data provided has resulted in exploring some SFDA (Source Free Domain Adaptation) approaches. Ultimately I want to realize SFDA as mentioned in <a href=\"https://research.google/blog/in-search-of-a-generalizable-method-for-source-free-domain-adaptation/\" target=\"_blank\">https://research.google/blog/in-search-of-a-generalizable-method-for-source-free-domain-adaptation/</a> but not able to do so.</p>\n<p>Need to use public LB score as CV. Not a good situation but local CV is quite high and there is evident domain shift in the test data<br>\nPhase 1:<br>\nPretraining using XCM data (<a href=\"https://huggingface.co/datasets/DBD-research-group/BirdSet)\" target=\"_blank\">https://huggingface.co/datasets/DBD-research-group/BirdSet)</a>, BirdClef 2021, 2022, 2023<br>\nTraining using Birdclef 2024 data. Training and augmentation as mentioned in <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66</a> is used<br>\nMax public/private score: 0.66 / 0.63</p>\n<p>Phase 2: Pseudo Labels<br>\nGeneral methodology used is in <a href=\"https://arxiv.org/pdf/2011.13439\" target=\"_blank\">https://arxiv.org/pdf/2011.13439</a> (UNSUPERVISED DOMAIN ADAPTATION FOR SPEECH RECOGNITION VIA UNCERTAINTY DRIVEN SELF-TRAINING)<br>\nstep 1) Generate pseudo labels in eval mode<br>\nstep 2) Generate pseudo label in train mode with dropout (noisy), generated 8 samples with different seed<br>\nstep 3) Calculate cosine distance of pairs (eval mode prediction, train mode prediction)<br>\nstep 4) Select ulabeled frames ( file + start) having cosine distance &lt; ( mean(cosine) - std(cosine))<br>\nIt selected around 21 K frames for unlabeled data</p>\n<p>`<br>\npreds = []<br>\npreds1=train_fold(True)<br>\npreds.append(preds1)<br>\nfor i in range(8):<br>\n    set_seed(i)<br>\n    preds2=train_fold(False)<br>\n    preds.append(preds2)<br>\npredsa=np.stack([preds],axis=3)</p>\n<p>diffs = []<br>\nfor i in range(predsa.shape[1]):<br>\n    p0 = predsa[0,i][:,0]<br>\n    darr = []<br>\n    for j in range(1,9):<br>\n        pp = predsa[j,i][:,0]<br>\n        d = cosine(p0,pp)<br>\n        darr.append(d)<br>\n    diffs.append(darr)<br>\ndiffs = np.array(diffs)<br>\nsoundscape_meta_df['cosine'] = diffs.sum(axis=1)<br>\nsoundscape_meta_df['preds'] = [ p for p in predsa[0]]<br>\n`</p>\n<p>Phase 3: Using filtered unlabeled data with pseudo labels and competition data<br>\nstep 1) Introduced augmentation by added noises generated from nocall + ESC 50. Training with noise now didn't hurt LB score which was earlier occurring only with competition data.<br>\nstep 2) Pretraining with noise augmentation<br>\nstep 3) Training with noise augmentation, trained 2 models with eca_nfnet_l0 and effnet_b0 backends<br>\nstep 4) Ensemble of two gives following public/private LB: 0.702/ 0.666</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2Fd451f34fd6be181772a5d38ea7208ccf%2Flb1.PNG?generation=1718092773801667&amp;alt=media\"></p>\n<p>Opportunity lost, I was thinking to create ensemble of some other 2 models, but driven by LB to select the maximum scoring ensemble<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F9e825a1e53c092c23867d04e67fd9bf4%2Flb2.PNG?generation=1718092978601455&amp;alt=media\"></p>\n<p>After reading discussions I should have used openvino to create better ensembles.. I was using pytorch script and limited to 2 models.</p>",
  "messages": [
    {
      "id": "2866230",
      "postDate": "06/11/2024 08:03:08",
      "content": "<p>Thanks to competition host for hosting this year BirdClef a bit different  as compared to previous competitions. Data provided has resulted in exploring some SFDA (Source Free Domain Adaptation) approaches. Ultimately I want to realize SFDA as mentioned in <a href=\"https://research.google/blog/in-search-of-a-generalizable-method-for-source-free-domain-adaptation/\" target=\"_blank\">https://research.google/blog/in-search-of-a-generalizable-method-for-source-free-domain-adaptation/</a> but not able to do so.</p>\n<p>Need to use public LB score as CV. Not a good situation but local CV is quite high and there is evident domain shift in the test data<br>\nPhase 1:<br>\nPretraining using XCM data (<a href=\"https://huggingface.co/datasets/DBD-research-group/BirdSet)\" target=\"_blank\">https://huggingface.co/datasets/DBD-research-group/BirdSet)</a>, BirdClef 2021, 2022, 2023<br>\nTraining using Birdclef 2024 data. Training and augmentation as mentioned in <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66</a> is used<br>\nMax public/private score: 0.66 / 0.63</p>\n<p>Phase 2: Pseudo Labels<br>\nGeneral methodology used is in <a href=\"https://arxiv.org/pdf/2011.13439\" target=\"_blank\">https://arxiv.org/pdf/2011.13439</a> (UNSUPERVISED DOMAIN ADAPTATION FOR SPEECH RECOGNITION VIA UNCERTAINTY DRIVEN SELF-TRAINING)<br>\nstep 1) Generate pseudo labels in eval mode<br>\nstep 2) Generate pseudo label in train mode with dropout (noisy), generated 8 samples with different seed<br>\nstep 3) Calculate cosine distance of pairs (eval mode prediction, train mode prediction)<br>\nstep 4) Select ulabeled frames ( file + start) having cosine distance &lt; ( mean(cosine) - std(cosine))<br>\nIt selected around 21 K frames for unlabeled data</p>\n<p>`<br>\npreds = []<br>\npreds1=train_fold(True)<br>\npreds.append(preds1)<br>\nfor i in range(8):<br>\n    set_seed(i)<br>\n    preds2=train_fold(False)<br>\n    preds.append(preds2)<br>\npredsa=np.stack([preds],axis=3)</p>\n<p>diffs = []<br>\nfor i in range(predsa.shape[1]):<br>\n    p0 = predsa[0,i][:,0]<br>\n    darr = []<br>\n    for j in range(1,9):<br>\n        pp = predsa[j,i][:,0]<br>\n        d = cosine(p0,pp)<br>\n        darr.append(d)<br>\n    diffs.append(darr)<br>\ndiffs = np.array(diffs)<br>\nsoundscape_meta_df['cosine'] = diffs.sum(axis=1)<br>\nsoundscape_meta_df['preds'] = [ p for p in predsa[0]]<br>\n`</p>\n<p>Phase 3: Using filtered unlabeled data with pseudo labels and competition data<br>\nstep 1) Introduced augmentation by added noises generated from nocall + ESC 50. Training with noise now didn't hurt LB score which was earlier occurring only with competition data.<br>\nstep 2) Pretraining with noise augmentation<br>\nstep 3) Training with noise augmentation, trained 2 models with eca_nfnet_l0 and effnet_b0 backends<br>\nstep 4) Ensemble of two gives following public/private LB: 0.702/ 0.666</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2Fd451f34fd6be181772a5d38ea7208ccf%2Flb1.PNG?generation=1718092773801667&amp;alt=media\"></p>\n<p>Opportunity lost, I was thinking to create ensemble of some other 2 models, but driven by LB to select the maximum scoring ensemble<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F9e825a1e53c092c23867d04e67fd9bf4%2Flb2.PNG?generation=1718092978601455&amp;alt=media\"></p>\n<p>After reading discussions I should have used openvino to create better ensembles.. I was using pytorch script and limited to 2 models.</p>",
      "rawMarkdown": "Thanks to competition host for hosting this year BirdClef a bit different  as compared to previous competitions. Data provided has resulted in exploring some SFDA (Source Free Domain Adaptation) approaches. Ultimately I want to realize SFDA as mentioned in https://research.google/blog/in-search-of-a-generalizable-method-for-source-free-domain-adaptation/ but not able to do so.\n\nNeed to use public LB score as CV. Not a good situation but local CV is quite high and there is evident domain shift in the test data\nPhase 1:\nPretraining using XCM data (https://huggingface.co/datasets/DBD-research-group/BirdSet), BirdClef 2021, 2022, 2023\nTraining using Birdclef 2024 data. Training and augmentation as mentioned in https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66 is used\nMax public/private score: 0.66 / 0.63\n\nPhase 2: Pseudo Labels\nGeneral methodology used is in https://arxiv.org/pdf/2011.13439 (UNSUPERVISED DOMAIN ADAPTATION FOR SPEECH RECOGNITION VIA UNCERTAINTY DRIVEN SELF-TRAINING)\nstep 1) Generate pseudo labels in eval mode\nstep 2) Generate pseudo label in train mode with dropout (noisy), generated 8 samples with different seed\nstep 3) Calculate cosine distance of pairs (eval mode prediction, train mode prediction)\nstep 4) Select ulabeled frames ( file + start) having cosine distance < ( mean(cosine) - std(cosine))\nIt selected around 21 K frames for unlabeled data\n\n`\npreds = []\npreds1=train_fold(True)\npreds.append(preds1)\nfor i in range(8):\n    set_seed(i)\n    preds2=train_fold(False)\n    preds.append(preds2)\npredsa=np.stack([preds],axis=3)\n\n\ndiffs = []\nfor i in range(predsa.shape[1]):\n    p0 = predsa[0,i][:,0]\n    darr = []\n    for j in range(1,9):\n        pp = predsa[j,i][:,0]\n        d = cosine(p0,pp)\n        darr.append(d)\n    diffs.append(darr)\ndiffs = np.array(diffs)\nsoundscape_meta_df['cosine'] = diffs.sum(axis=1)\nsoundscape_meta_df['preds'] = [ p for p in predsa[0]]\n`\n\nPhase 3: Using filtered unlabeled data with pseudo labels and competition data\nstep 1) Introduced augmentation by added noises generated from nocall + ESC 50. Training with noise now didn't hurt LB score which was earlier occurring only with competition data.\nstep 2) Pretraining with noise augmentation\nstep 3) Training with noise augmentation, trained 2 models with eca_nfnet_l0 and effnet_b0 backends\nstep 4) Ensemble of two gives following public/private LB: 0.702/ 0.666\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2Fd451f34fd6be181772a5d38ea7208ccf%2Flb1.PNG?generation=1718092773801667&alt=media)\n\nOpportunity lost, I was thinking to create ensemble of some other 2 models, but driven by LB to select the maximum scoring ensemble\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F9e825a1e53c092c23867d04e67fd9bf4%2Flb2.PNG?generation=1718092978601455&alt=media)\n\nAfter reading discussions I should have used openvino to create better ensembles.. I was using pytorch script and limited to 2 models.",
      "votes": null
    },
    {
      "id": "2866304",
      "postDate": "06/11/2024 08:39:28",
      "content": "<p>Really good job!</p>",
      "rawMarkdown": "Really good job!",
      "votes": null
    },
    {
      "id": "2878312",
      "postDate": "06/18/2024 21:18:53",
      "content": "<p>Great to see that someone tried NOTELA! Thanks for the write-up.</p>",
      "rawMarkdown": "Great to see that someone tried NOTELA! Thanks for the write-up.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2866304,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "06/11/2024 08:39:28",
      "content": "<p>Really good job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2878312,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "06/18/2024 21:18:53",
      "content": "<p>Great to see that someone tried NOTELA! Thanks for the write-up.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2866230": "Thanks to competition host for hosting this year BirdClef a bit different  as compared to previous competitions. Data provided has resulted in exploring some SFDA (Source Free Domain Adaptation) approaches. Ultimately I want to realize SFDA as mentioned in https://research.google/blog/in-search-of-a-generalizable-method-for-source-free-domain-adaptation/ but not able to do so.\n\nNeed to use public LB score as CV. Not a good situation but local CV is quite high and there is evident domain shift in the test data\nPhase 1:\nPretraining using XCM data (https://huggingface.co/datasets/DBD-research-group/BirdSet), BirdClef 2021, 2022, 2023\nTraining using Birdclef 2024 data. Training and augmentation as mentioned in https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66 is used\nMax public/private score: 0.66 / 0.63\n\nPhase 2: Pseudo Labels\nGeneral methodology used is in https://arxiv.org/pdf/2011.13439 (UNSUPERVISED DOMAIN ADAPTATION FOR SPEECH RECOGNITION VIA UNCERTAINTY DRIVEN SELF-TRAINING)\nstep 1) Generate pseudo labels in eval mode\nstep 2) Generate pseudo label in train mode with dropout (noisy), generated 8 samples with different seed\nstep 3) Calculate cosine distance of pairs (eval mode prediction, train mode prediction)\nstep 4) Select ulabeled frames ( file + start) having cosine distance < ( mean(cosine) - std(cosine))\nIt selected around 21 K frames for unlabeled data\n\n`\npreds = []\npreds1=train_fold(True)\npreds.append(preds1)\nfor i in range(8):\n    set_seed(i)\n    preds2=train_fold(False)\n    preds.append(preds2)\npredsa=np.stack([preds],axis=3)\n\n\ndiffs = []\nfor i in range(predsa.shape[1]):\n    p0 = predsa[0,i][:,0]\n    darr = []\n    for j in range(1,9):\n        pp = predsa[j,i][:,0]\n        d = cosine(p0,pp)\n        darr.append(d)\n    diffs.append(darr)\ndiffs = np.array(diffs)\nsoundscape_meta_df['cosine'] = diffs.sum(axis=1)\nsoundscape_meta_df['preds'] = [ p for p in predsa[0]]\n`\n\nPhase 3: Using filtered unlabeled data with pseudo labels and competition data\nstep 1) Introduced augmentation by added noises generated from nocall + ESC 50. Training with noise now didn't hurt LB score which was earlier occurring only with competition data.\nstep 2) Pretraining with noise augmentation\nstep 3) Training with noise augmentation, trained 2 models with eca_nfnet_l0 and effnet_b0 backends\nstep 4) Ensemble of two gives following public/private LB: 0.702/ 0.666\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2Fd451f34fd6be181772a5d38ea7208ccf%2Flb1.PNG?generation=1718092773801667&alt=media)\n\nOpportunity lost, I was thinking to create ensemble of some other 2 models, but driven by LB to select the maximum scoring ensemble\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F9e825a1e53c092c23867d04e67fd9bf4%2Flb2.PNG?generation=1718092978601455&alt=media)\n\nAfter reading discussions I should have used openvino to create better ensembles.. I was using pytorch script and limited to 2 models.",
    "2866304": "Really good job!",
    "2878312": "Great to see that someone tried NOTELA! Thanks for the write-up."
  },
  "source": "meta"
}