{
  "id": 498582,
  "title": "Analyzing train vs test data using spectrogram statistics",
  "url": "/competitions/birdclef-2024/discussion/498582",
  "author_name": "",
  "post_date": "2024-04-28T22:27:01.998568100Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Like <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/493317\" target=\"_blank\">many</a> <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498404\" target=\"_blank\">others</a>, I'm struggling to correlate my own local cross validation scores (CV) with the leaderboard score (LB). I first looked for <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/494134\" target=\"_blank\">duplicate files</a> to remove any training/validation leakage, but only ~150 duplicate files have been found so far out of ~24,500 in the dataset.</p>\n<p>Now I'm trying to understand the difference/shift between the training data and the unlabeled soundscapes, since the model may have a hard time applying knowledge learned from one dataset to the other if the difference between them is large. (The unlabeled soundscapes are from the same locations as the test data, so they should be representative.)</p>\n<p>Aggregating frequency data from each dataset is one way to quantify the shift. <a href=\"https://www.kaggle.com/code/robbynevels/bc24-test-train-shift-in-spectrograms\" target=\"_blank\">Check out this notebook</a> to play with different ways of generating frequency statistics.</p>\n<p>Here is the mean and standard deviation of the decibels of each mel filterbank from the first 5 seconds of each file:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5737610%2Fcd336e0b2c15443c9ed3b62b1139489b%2FScreenshot%202024-04-28%20at%205.24.30PM.png?generation=1714343096143319&amp;alt=media\" alt=\"chart of mel filterbank value statistics\"></p>\n<p>This shows quite a difference across the spectrum. I'm not sure how to interpret it or how to correct for it during training and evaluation. Does anyone have ideas? Or have you developed another way to analyze this shift?</p>",
  "messages": [
    {
      "id": "2781634",
      "postDate": "04/28/2024 22:27:02",
      "content": "<p>Like <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/493317\" target=\"_blank\">many</a> <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/498404\" target=\"_blank\">others</a>, I'm struggling to correlate my own local cross validation scores (CV) with the leaderboard score (LB). I first looked for <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/494134\" target=\"_blank\">duplicate files</a> to remove any training/validation leakage, but only ~150 duplicate files have been found so far out of ~24,500 in the dataset.</p>\n<p>Now I'm trying to understand the difference/shift between the training data and the unlabeled soundscapes, since the model may have a hard time applying knowledge learned from one dataset to the other if the difference between them is large. (The unlabeled soundscapes are from the same locations as the test data, so they should be representative.)</p>\n<p>Aggregating frequency data from each dataset is one way to quantify the shift. <a href=\"https://www.kaggle.com/code/robbynevels/bc24-test-train-shift-in-spectrograms\" target=\"_blank\">Check out this notebook</a> to play with different ways of generating frequency statistics.</p>\n<p>Here is the mean and standard deviation of the decibels of each mel filterbank from the first 5 seconds of each file:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5737610%2Fcd336e0b2c15443c9ed3b62b1139489b%2FScreenshot%202024-04-28%20at%205.24.30PM.png?generation=1714343096143319&amp;alt=media\" alt=\"chart of mel filterbank value statistics\"></p>\n<p>This shows quite a difference across the spectrum. I'm not sure how to interpret it or how to correct for it during training and evaluation. Does anyone have ideas? Or have you developed another way to analyze this shift?</p>",
      "rawMarkdown": "Like [many](https://www.kaggle.com/competitions/birdclef-2024/discussion/493317) [others](https://www.kaggle.com/competitions/birdclef-2024/discussion/498404), I'm struggling to correlate my own local cross validation scores (CV) with the leaderboard score (LB). I first looked for [duplicate files](https://www.kaggle.com/competitions/birdclef-2024/discussion/494134) to remove any training/validation leakage, but only ~150 duplicate files have been found so far out of ~24,500 in the dataset.\n\nNow I'm trying to understand the difference/shift between the training data and the unlabeled soundscapes, since the model may have a hard time applying knowledge learned from one dataset to the other if the difference between them is large. (The unlabeled soundscapes are from the same locations as the test data, so they should be representative.)\n\nAggregating frequency data from each dataset is one way to quantify the shift. [Check out this notebook](https://www.kaggle.com/code/robbynevels/bc24-test-train-shift-in-spectrograms) to play with different ways of generating frequency statistics.\n\nHere is the mean and standard deviation of the decibels of each mel filterbank from the first 5 seconds of each file:\n\n![chart of mel filterbank value statistics](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5737610%2Fcd336e0b2c15443c9ed3b62b1139489b%2FScreenshot%202024-04-28%20at%205.24.30PM.png?generation=1714343096143319&alt=media)\n\nThis shows quite a difference across the spectrum. I'm not sure how to interpret it or how to correct for it during training and evaluation. Does anyone have ideas? Or have you developed another way to analyze this shift?",
      "votes": null
    },
    {
      "id": "2781986",
      "postDate": "04/29/2024 05:00:16",
      "content": "<p>Interestingly, unmarked signals have louder high frequencies.</p>",
      "rawMarkdown": "Interestingly, unmarked signals have louder high frequencies.",
      "votes": null
    },
    {
      "id": "2788473",
      "postDate": "05/02/2024 08:48:29",
      "content": "<p>Did handling the duplicates change your local CV significantly?</p>",
      "rawMarkdown": "Did handling the duplicates change your local CV significantly?",
      "votes": null
    },
    {
      "id": "2789203",
      "postDate": "05/02/2024 15:39:38",
      "content": "<p>It made a small CV improvement for me, but it was so small that I’m not sure it was statistically significant. I didn’t see an impact on LB.</p>",
      "rawMarkdown": "It made a small CV improvement for me, but it was so small that I’m not sure it was statistically significant. I didn’t see an impact on LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2781986,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "04/29/2024 05:00:16",
      "content": "<p>Interestingly, unmarked signals have louder high frequencies.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2788473,
      "author_name": "hugodeheer",
      "author_url": "",
      "post_date": "05/02/2024 08:48:29",
      "content": "<p>Did handling the duplicates change your local CV significantly?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2789203,
          "author_name": "robbynevels",
          "author_url": "",
          "post_date": "05/02/2024 15:39:38",
          "content": "<p>It made a small CV improvement for me, but it was so small that I’m not sure it was statistically significant. I didn’t see an impact on LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2781634": "Like [many](https://www.kaggle.com/competitions/birdclef-2024/discussion/493317) [others](https://www.kaggle.com/competitions/birdclef-2024/discussion/498404), I'm struggling to correlate my own local cross validation scores (CV) with the leaderboard score (LB). I first looked for [duplicate files](https://www.kaggle.com/competitions/birdclef-2024/discussion/494134) to remove any training/validation leakage, but only ~150 duplicate files have been found so far out of ~24,500 in the dataset.\n\nNow I'm trying to understand the difference/shift between the training data and the unlabeled soundscapes, since the model may have a hard time applying knowledge learned from one dataset to the other if the difference between them is large. (The unlabeled soundscapes are from the same locations as the test data, so they should be representative.)\n\nAggregating frequency data from each dataset is one way to quantify the shift. [Check out this notebook](https://www.kaggle.com/code/robbynevels/bc24-test-train-shift-in-spectrograms) to play with different ways of generating frequency statistics.\n\nHere is the mean and standard deviation of the decibels of each mel filterbank from the first 5 seconds of each file:\n\n![chart of mel filterbank value statistics](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5737610%2Fcd336e0b2c15443c9ed3b62b1139489b%2FScreenshot%202024-04-28%20at%205.24.30PM.png?generation=1714343096143319&alt=media)\n\nThis shows quite a difference across the spectrum. I'm not sure how to interpret it or how to correct for it during training and evaluation. Does anyone have ideas? Or have you developed another way to analyze this shift?",
    "2781986": "Interestingly, unmarked signals have louder high frequencies.",
    "2788473": "Did handling the duplicates change your local CV significantly?",
    "2789203": "It made a small CV improvement for me, but it was so small that I’m not sure it was statistically significant. I didn’t see an impact on LB."
  },
  "source": "meta"
}