{
  "id": 496956,
  "title": "BirdTrax: Labeled soundscape generator",
  "url": "/competitions/birdclef-2024/discussion/496956",
  "author_name": "",
  "post_date": "2024-04-23T04:41:27.355976400Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I built a notebook that generates 4-minute labeled bird-call soundscapes:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdtrax-birdclef-2024-labeled-soundscapes\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdtrax-birdclef-2024-labeled-soundscapes</a></p>\n<p>It works by overlapping random 5-second segments of the labeled training data.</p>\n<p>You can configure it to overlap as many different species as you like.  There are a bunch of settings to play with…</p>\n<p>It only uses recordings with 4+ quality ratings and don't have secondary labels.</p>\n<p>By default 3x samples are overlapped for -each- species present at any point in the soundscape.  I did this to reduce the problem with some 5-second segments in the train data not containing calls for the labeled species.</p>\n<p>The notebook outputs the soundscapes as OGGs and the labeled data to a CSV file intended to mirror the format of submission.csv</p>\n<p>You might be able to combine this with the logic in <a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">https://www.kaggle.com/code/metric/birdclef-roc-auc</a> to actually score your model…  (maybe I'll make a demo notebook for that…)</p>\n<p>The soundscape generator by default does a train / validate split (and uses the validate files for the bird-calls).  If you wanted to get really fancy - you could train on part of the data - and then validate against soundscapes built from the other part.</p>\n<p>The saved notebook's output has 10 labeled OGG files + corresponding CSV if you want to play.</p>\n<p>At some point I'll export a larger dataset of labeled soundscapes.</p>\n<p>Hope this is useful to someone.</p>\n<p>-Rich</p>",
  "messages": [
    {
      "id": "2768886",
      "postDate": "04/23/2024 04:41:27",
      "content": "<p>I built a notebook that generates 4-minute labeled bird-call soundscapes:</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdtrax-birdclef-2024-labeled-soundscapes\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdtrax-birdclef-2024-labeled-soundscapes</a></p>\n<p>It works by overlapping random 5-second segments of the labeled training data.</p>\n<p>You can configure it to overlap as many different species as you like.  There are a bunch of settings to play with…</p>\n<p>It only uses recordings with 4+ quality ratings and don't have secondary labels.</p>\n<p>By default 3x samples are overlapped for -each- species present at any point in the soundscape.  I did this to reduce the problem with some 5-second segments in the train data not containing calls for the labeled species.</p>\n<p>The notebook outputs the soundscapes as OGGs and the labeled data to a CSV file intended to mirror the format of submission.csv</p>\n<p>You might be able to combine this with the logic in <a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">https://www.kaggle.com/code/metric/birdclef-roc-auc</a> to actually score your model…  (maybe I'll make a demo notebook for that…)</p>\n<p>The soundscape generator by default does a train / validate split (and uses the validate files for the bird-calls).  If you wanted to get really fancy - you could train on part of the data - and then validate against soundscapes built from the other part.</p>\n<p>The saved notebook's output has 10 labeled OGG files + corresponding CSV if you want to play.</p>\n<p>At some point I'll export a larger dataset of labeled soundscapes.</p>\n<p>Hope this is useful to someone.</p>\n<p>-Rich</p>",
      "rawMarkdown": "I built a notebook that generates 4-minute labeled bird-call soundscapes:\n\nhttps://www.kaggle.com/code/richolson/birdtrax-birdclef-2024-labeled-soundscapes\n\nIt works by overlapping random 5-second segments of the labeled training data.\n\nYou can configure it to overlap as many different species as you like.  There are a bunch of settings to play with...\n\nIt only uses recordings with 4+ quality ratings and don't have secondary labels.\n\nBy default 3x samples are overlapped for -each- species present at any point in the soundscape.  I did this to reduce the problem with some 5-second segments in the train data not containing calls for the labeled species.\n\nThe notebook outputs the soundscapes as OGGs and the labeled data to a CSV file intended to mirror the format of submission.csv\n\nYou might be able to combine this with the logic in https://www.kaggle.com/code/metric/birdclef-roc-auc to actually score your model...  (maybe I'll make a demo notebook for that...)\n\nThe soundscape generator by default does a train / validate split (and uses the validate files for the bird-calls).  If you wanted to get really fancy - you could train on part of the data - and then validate against soundscapes built from the other part.\n\nThe saved notebook's output has 10 labeled OGG files + corresponding CSV if you want to play.\n\nAt some point I'll export a larger dataset of labeled soundscapes.\n\nHope this is useful to someone.\n\n-Rich",
      "votes": null
    },
    {
      "id": "2770362",
      "postDate": "04/23/2024 19:56:16",
      "content": "<p>Is it possible to approach the LB score on hidden testset using your soundscapes?</p>",
      "rawMarkdown": "Is it possible to approach the LB score on hidden testset using your soundscapes?",
      "votes": null
    },
    {
      "id": "2770414",
      "postDate": "04/23/2024 20:30:11",
      "content": "<p>that's a very good question!  I haven't tested yet…</p>\n<p>I don't know how closely these soundscapes reflect the test data - but I figure it's at least something in the same format.</p>\n<p>one difference is that these soundscapes use all the train data - while the test data is only on the Western Ghats.  Would be fairly easy to make the soundscape generated from only Western Ghat's train data (maybe 15% of train data) - but if you then filtered on quality - the sample set might start getting kind of small.</p>\n<p>I have a public .61 LB notebook that I will probably adapt to score against the generated soundscapes.</p>",
      "rawMarkdown": "that's a very good question!  I haven't tested yet...\n\nI don't know how closely these soundscapes reflect the test data - but I figure it's at least something in the same format.\n\none difference is that these soundscapes use all the train data - while the test data is only on the Western Ghats.  Would be fairly easy to make the soundscape generated from only Western Ghat's train data (maybe 15% of train data) - but if you then filtered on quality - the sample set might start getting kind of small.\n\nI have a public .61 LB notebook that I will probably adapt to score against the generated soundscapes.",
      "votes": null
    },
    {
      "id": "2770495",
      "postDate": "04/23/2024 21:28:30",
      "content": "<p>My understanding is that the host is able to choose which columns get scored by passing the \"solution\" dataframe to the score function. This information is not available but could be responsible for the mismatch between LB and local score, as long as the model does not recognize all classes equally well.</p>",
      "rawMarkdown": "My understanding is that the host is able to choose which columns get scored by passing the \"solution\" dataframe to the score function. This information is not available but could be responsible for the mismatch between LB and local score, as long as the model does not recognize all classes equally well.",
      "votes": null
    },
    {
      "id": "2770597",
      "postDate": "04/23/2024 23:22:29",
      "content": "<p>my understanding is that any columns that have species present will be scored.  (those that don't - don' t matter)</p>",
      "rawMarkdown": "my understanding is that any columns that have species present will be scored.  (those that don't - don' t matter)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2770362,
      "author_name": "vi2018",
      "author_url": "",
      "post_date": "04/23/2024 19:56:16",
      "content": "<p>Is it possible to approach the LB score on hidden testset using your soundscapes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2770414,
          "author_name": "richolson",
          "author_url": "",
          "post_date": "04/23/2024 20:30:11",
          "content": "<p>that's a very good question!  I haven't tested yet…</p>\n<p>I don't know how closely these soundscapes reflect the test data - but I figure it's at least something in the same format.</p>\n<p>one difference is that these soundscapes use all the train data - while the test data is only on the Western Ghats.  Would be fairly easy to make the soundscape generated from only Western Ghat's train data (maybe 15% of train data) - but if you then filtered on quality - the sample set might start getting kind of small.</p>\n<p>I have a public .61 LB notebook that I will probably adapt to score against the generated soundscapes.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2770495,
              "author_name": "vi2018",
              "author_url": "",
              "post_date": "04/23/2024 21:28:30",
              "content": "<p>My understanding is that the host is able to choose which columns get scored by passing the \"solution\" dataframe to the score function. This information is not available but could be responsible for the mismatch between LB and local score, as long as the model does not recognize all classes equally well.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2770597,
                  "author_name": "richolson",
                  "author_url": "",
                  "post_date": "04/23/2024 23:22:29",
                  "content": "<p>my understanding is that any columns that have species present will be scored.  (those that don't - don' t matter)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2768886": "I built a notebook that generates 4-minute labeled bird-call soundscapes:\n\nhttps://www.kaggle.com/code/richolson/birdtrax-birdclef-2024-labeled-soundscapes\n\nIt works by overlapping random 5-second segments of the labeled training data.\n\nYou can configure it to overlap as many different species as you like.  There are a bunch of settings to play with...\n\nIt only uses recordings with 4+ quality ratings and don't have secondary labels.\n\nBy default 3x samples are overlapped for -each- species present at any point in the soundscape.  I did this to reduce the problem with some 5-second segments in the train data not containing calls for the labeled species.\n\nThe notebook outputs the soundscapes as OGGs and the labeled data to a CSV file intended to mirror the format of submission.csv\n\nYou might be able to combine this with the logic in https://www.kaggle.com/code/metric/birdclef-roc-auc to actually score your model...  (maybe I'll make a demo notebook for that...)\n\nThe soundscape generator by default does a train / validate split (and uses the validate files for the bird-calls).  If you wanted to get really fancy - you could train on part of the data - and then validate against soundscapes built from the other part.\n\nThe saved notebook's output has 10 labeled OGG files + corresponding CSV if you want to play.\n\nAt some point I'll export a larger dataset of labeled soundscapes.\n\nHope this is useful to someone.\n\n-Rich",
    "2770362": "Is it possible to approach the LB score on hidden testset using your soundscapes?",
    "2770414": "that's a very good question!  I haven't tested yet...\n\nI don't know how closely these soundscapes reflect the test data - but I figure it's at least something in the same format.\n\none difference is that these soundscapes use all the train data - while the test data is only on the Western Ghats.  Would be fairly easy to make the soundscape generated from only Western Ghat's train data (maybe 15% of train data) - but if you then filtered on quality - the sample set might start getting kind of small.\n\nI have a public .61 LB notebook that I will probably adapt to score against the generated soundscapes.",
    "2770495": "My understanding is that the host is able to choose which columns get scored by passing the \"solution\" dataframe to the score function. This information is not available but could be responsible for the mismatch between LB and local score, as long as the model does not recognize all classes equally well.",
    "2770597": "my understanding is that any columns that have species present will be scored.  (those that don't - don' t matter)"
  },
  "source": "meta"
}