{
  "id": 494134,
  "title": "Duplicate Files",
  "url": "/competitions/birdclef-2024/discussion/494134",
  "author_name": "",
  "post_date": "2024-04-16T05:28:53.745319800Z",
  "votes": 38,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I found 148 pairs of duplicate audio files in the dataset, listed &amp; displayed <a href=\"https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/\" target=\"_blank\">here</a>. Duplicate files are problematic because they can cause overlap between training &amp; validation data, leading to overfitting. I will update that notebook as more duplicates are found, so please let me know if you find others or see a mistake.</p>\n<p>Finding duplicate files is tricky because they're not always exact copies -- there can be noise, recompression artifacts, or cropping. So instead of directly comparing audio or spectrograms, I compared embeddings from the <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier/4\" target=\"_blank\">google bird vocalization classifier</a>, since embeddings are a higher level representation of the data and should be robust to non-semantic changes. Here's <a href=\"https://www.kaggle.com/code/robbynevels/bc24-data-identifying-duplicates-w-embedddings\" target=\"_blank\">a very messy notebook</a> showing the method I used, applying thresholds on embedding distances and sound volume from the first and last 5 seconds of each recording. (and here's my notebook from <a href=\"https://www.kaggle.com/code/robbynevels/identifying-duplicates-with-embeddings\" target=\"_blank\">last year</a> doing something similar)</p>\n<p>I also made my <a href=\"https://www.kaggle.com/code/robbynevels/bc24-google-bird-model-embeddings-predict-score/notebook\" target=\"_blank\">google bird embeddings notebook public</a>, so feel free to use the embeddings &amp; model predictions for other analysis.</p>\n<p>By the way, I’ve found that some of these duplicate pairs have different primary labels, like \"purher1/XC467373.ogg\" and \"graher1/XC467373.ogg\" for example. Maybe the competition hosts or someone more familiar with Xeno-Canto might know why that would happen? </p>",
  "messages": [
    {
      "id": "2754567",
      "postDate": "04/16/2024 05:28:53",
      "content": "<p>I found 148 pairs of duplicate audio files in the dataset, listed &amp; displayed <a href=\"https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/\" target=\"_blank\">here</a>. Duplicate files are problematic because they can cause overlap between training &amp; validation data, leading to overfitting. I will update that notebook as more duplicates are found, so please let me know if you find others or see a mistake.</p>\n<p>Finding duplicate files is tricky because they're not always exact copies -- there can be noise, recompression artifacts, or cropping. So instead of directly comparing audio or spectrograms, I compared embeddings from the <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier/4\" target=\"_blank\">google bird vocalization classifier</a>, since embeddings are a higher level representation of the data and should be robust to non-semantic changes. Here's <a href=\"https://www.kaggle.com/code/robbynevels/bc24-data-identifying-duplicates-w-embedddings\" target=\"_blank\">a very messy notebook</a> showing the method I used, applying thresholds on embedding distances and sound volume from the first and last 5 seconds of each recording. (and here's my notebook from <a href=\"https://www.kaggle.com/code/robbynevels/identifying-duplicates-with-embeddings\" target=\"_blank\">last year</a> doing something similar)</p>\n<p>I also made my <a href=\"https://www.kaggle.com/code/robbynevels/bc24-google-bird-model-embeddings-predict-score/notebook\" target=\"_blank\">google bird embeddings notebook public</a>, so feel free to use the embeddings &amp; model predictions for other analysis.</p>\n<p>By the way, I’ve found that some of these duplicate pairs have different primary labels, like \"purher1/XC467373.ogg\" and \"graher1/XC467373.ogg\" for example. Maybe the competition hosts or someone more familiar with Xeno-Canto might know why that would happen? </p>",
      "rawMarkdown": "I found 148 pairs of duplicate audio files in the dataset, listed & displayed [here](https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/). Duplicate files are problematic because they can cause overlap between training & validation data, leading to overfitting. I will update that notebook as more duplicates are found, so please let me know if you find others or see a mistake.\n\nFinding duplicate files is tricky because they're not always exact copies -- there can be noise, recompression artifacts, or cropping. So instead of directly comparing audio or spectrograms, I compared embeddings from the [google bird vocalization classifier](https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier/4), since embeddings are a higher level representation of the data and should be robust to non-semantic changes. Here's [a very messy notebook](https://www.kaggle.com/code/robbynevels/bc24-data-identifying-duplicates-w-embedddings) showing the method I used, applying thresholds on embedding distances and sound volume from the first and last 5 seconds of each recording. (and here's my notebook from [last year](https://www.kaggle.com/code/robbynevels/identifying-duplicates-with-embeddings) doing something similar)\n\nI also made my [google bird embeddings notebook public](https://www.kaggle.com/code/robbynevels/bc24-google-bird-model-embeddings-predict-score/notebook), so feel free to use the embeddings & model predictions for other analysis.\n\nBy the way, I’ve found that some of these duplicate pairs have different primary labels, like \"purher1/XC467373.ogg\" and \"graher1/XC467373.ogg\" for example. Maybe the competition hosts or someone more familiar with Xeno-Canto might know why that would happen?",
      "votes": null
    },
    {
      "id": "2756020",
      "postDate": "04/16/2024 19:31:13",
      "content": "<p>Nice find!  Thank you for sharing.</p>",
      "rawMarkdown": "Nice find!  Thank you for sharing.",
      "votes": null
    },
    {
      "id": "2756142",
      "postDate": "04/16/2024 22:27:42",
      "content": "<p>Looked at train_metadata.csv a little bit…</p>\n<p>The couple duplicates with different species labels I looked at manually did -not- have secondary labels. </p>\n<p>some of them at least had related species:</p>\n<p>woosan/XC825766.ogg Wood Sandpiper<br>\ngrnsan/XC825765.ogg Green Sandpiper</p>\n<p>whbwoo2/XC239509.ogg White-bellied Woodpecker<br>\nrufwoo2/XC239509.ogg Rufous Woodpecker</p>\n<p>Some of the dupes have different locations??? (lat / long)<br>\n\"zitcis1/XC302781.ogg\" = 12.2714     109.1138    <br>\n\"zitcis1/XC303866.ogg\" = 12.4419     109.2743</p>\n<p>So - looking at the URLs - the listings show different sub-species?:<br>\n<a href=\"https://xeno-canto.org/302781\" target=\"_blank\">https://xeno-canto.org/302781</a> Cisticola juncidis tinnabulans<br>\n<a href=\"https://xeno-canto.org/303866\" target=\"_blank\">https://xeno-canto.org/303866</a>  Cisticola juncidis</p>\n<p>I'll also note the recording quality on those last 2 was pretty bad (major background noise) - yet rated a 3?  Those 2 files clearly had the same audio - but were not byte-for-byte the same.</p>\n<p>I'm imagining someone correcting a misidentified species / moving stuff around (but leaving the original on accident)</p>",
      "rawMarkdown": "Looked at train_metadata.csv a little bit...\n\nThe couple duplicates with different species labels I looked at manually did -not- have secondary labels. \n\nsome of them at least had related species:\n\nwoosan/XC825766.ogg Wood Sandpiper\ngrnsan/XC825765.ogg Green Sandpiper\n\nwhbwoo2/XC239509.ogg White-bellied Woodpecker\nrufwoo2/XC239509.ogg Rufous Woodpecker\n\nSome of the dupes have different locations??? (lat / long)\n\"zitcis1/XC302781.ogg\" = 12.2714 \t109.1138 \t\n\"zitcis1/XC303866.ogg\" = 12.4419 \t109.2743\n\nSo - looking at the URLs - the listings show different sub-species?:\nhttps://xeno-canto.org/302781 Cisticola juncidis tinnabulans\nhttps://xeno-canto.org/303866  Cisticola juncidis\n\nI'll also note the recording quality on those last 2 was pretty bad (major background noise) - yet rated a 3?  Those 2 files clearly had the same audio - but were not byte-for-byte the same.\n\nI'm imagining someone correcting a misidentified species / moving stuff around (but leaving the original on accident)",
      "votes": null
    },
    {
      "id": "2756677",
      "postDate": "04/17/2024 06:09:36",
      "content": "<p>Great investigation folks, thanks for the detective work! My bet would be that the (near) duplicates are partially remaining artifacts from taxonomy changes and (as you suspected) corrected species ID. Unfortunately, these are super hard to detect, and I believe they should be in the original dataset because this is the kind of data biologists have to deal with. There is no actual ground truth - sometimes there are inconsistencies. If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.</p>\n<p>In terms of rating, these are also not super reliable, because they're based on user feedback (which is noisy at times).</p>",
      "rawMarkdown": "Great investigation folks, thanks for the detective work! My bet would be that the (near) duplicates are partially remaining artifacts from taxonomy changes and (as you suspected) corrected species ID. Unfortunately, these are super hard to detect, and I believe they should be in the original dataset because this is the kind of data biologists have to deal with. There is no actual ground truth - sometimes there are inconsistencies. If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.\n\nIn terms of rating, these are also not super reliable, because they're based on user feedback (which is noisy at times).",
      "votes": null
    },
    {
      "id": "2759346",
      "postDate": "04/18/2024 16:22:31",
      "content": "<blockquote>\n  <p>If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.</p>\n</blockquote>\n<p>Thresholding based on volume &amp; embedding distance appears to be fairly robust, but the thresholds have to be tuned whenever changing the embedding model, and human curation afterwards is still required to find occasional false positives. Also I'm not sure how many other duplicates there are that this technique is not detecting.</p>\n<p>I'd be interested in exploring this further. Let me know if you have any ideas on directions to try :)</p>",
      "rawMarkdown": "> If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.\n\nThresholding based on volume & embedding distance appears to be fairly robust, but the thresholds have to be tuned whenever changing the embedding model, and human curation afterwards is still required to find occasional false positives. Also I'm not sure how many other duplicates there are that this technique is not detecting.\n\nI'd be interested in exploring this further. Let me know if you have any ideas on directions to try :)",
      "votes": null
    },
    {
      "id": "2766563",
      "postDate": "04/21/2024 19:10:44",
      "content": "<p>Nice job! Thank you!</p>",
      "rawMarkdown": "Nice job! Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2756020,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "04/16/2024 19:31:13",
      "content": "<p>Nice find!  Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2756142,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "04/16/2024 22:27:42",
      "content": "<p>Looked at train_metadata.csv a little bit…</p>\n<p>The couple duplicates with different species labels I looked at manually did -not- have secondary labels. </p>\n<p>some of them at least had related species:</p>\n<p>woosan/XC825766.ogg Wood Sandpiper<br>\ngrnsan/XC825765.ogg Green Sandpiper</p>\n<p>whbwoo2/XC239509.ogg White-bellied Woodpecker<br>\nrufwoo2/XC239509.ogg Rufous Woodpecker</p>\n<p>Some of the dupes have different locations??? (lat / long)<br>\n\"zitcis1/XC302781.ogg\" = 12.2714     109.1138    <br>\n\"zitcis1/XC303866.ogg\" = 12.4419     109.2743</p>\n<p>So - looking at the URLs - the listings show different sub-species?:<br>\n<a href=\"https://xeno-canto.org/302781\" target=\"_blank\">https://xeno-canto.org/302781</a> Cisticola juncidis tinnabulans<br>\n<a href=\"https://xeno-canto.org/303866\" target=\"_blank\">https://xeno-canto.org/303866</a>  Cisticola juncidis</p>\n<p>I'll also note the recording quality on those last 2 was pretty bad (major background noise) - yet rated a 3?  Those 2 files clearly had the same audio - but were not byte-for-byte the same.</p>\n<p>I'm imagining someone correcting a misidentified species / moving stuff around (but leaving the original on accident)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2756677,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "04/17/2024 06:09:36",
      "content": "<p>Great investigation folks, thanks for the detective work! My bet would be that the (near) duplicates are partially remaining artifacts from taxonomy changes and (as you suspected) corrected species ID. Unfortunately, these are super hard to detect, and I believe they should be in the original dataset because this is the kind of data biologists have to deal with. There is no actual ground truth - sometimes there are inconsistencies. If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.</p>\n<p>In terms of rating, these are also not super reliable, because they're based on user feedback (which is noisy at times).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2759346,
          "author_name": "robbynevels",
          "author_url": "",
          "post_date": "04/18/2024 16:22:31",
          "content": "<blockquote>\n  <p>If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.</p>\n</blockquote>\n<p>Thresholding based on volume &amp; embedding distance appears to be fairly robust, but the thresholds have to be tuned whenever changing the embedding model, and human curation afterwards is still required to find occasional false positives. Also I'm not sure how many other duplicates there are that this technique is not detecting.</p>\n<p>I'd be interested in exploring this further. Let me know if you have any ideas on directions to try :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2766563,
      "author_name": "tjamali",
      "author_url": "",
      "post_date": "04/21/2024 19:10:44",
      "content": "<p>Nice job! Thank you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2754567": "I found 148 pairs of duplicate audio files in the dataset, listed & displayed [here](https://www.kaggle.com/code/robbynevels/bc24-duplicate-audio-files/). Duplicate files are problematic because they can cause overlap between training & validation data, leading to overfitting. I will update that notebook as more duplicates are found, so please let me know if you find others or see a mistake.\n\nFinding duplicate files is tricky because they're not always exact copies -- there can be noise, recompression artifacts, or cropping. So instead of directly comparing audio or spectrograms, I compared embeddings from the [google bird vocalization classifier](https://www.kaggle.com/models/google/bird-vocalization-classifier/TensorFlow2/bird-vocalization-classifier/4), since embeddings are a higher level representation of the data and should be robust to non-semantic changes. Here's [a very messy notebook](https://www.kaggle.com/code/robbynevels/bc24-data-identifying-duplicates-w-embedddings) showing the method I used, applying thresholds on embedding distances and sound volume from the first and last 5 seconds of each recording. (and here's my notebook from [last year](https://www.kaggle.com/code/robbynevels/identifying-duplicates-with-embeddings) doing something similar)\n\nI also made my [google bird embeddings notebook public](https://www.kaggle.com/code/robbynevels/bc24-google-bird-model-embeddings-predict-score/notebook), so feel free to use the embeddings & model predictions for other analysis.\n\nBy the way, I’ve found that some of these duplicate pairs have different primary labels, like \"purher1/XC467373.ogg\" and \"graher1/XC467373.ogg\" for example. Maybe the competition hosts or someone more familiar with Xeno-Canto might know why that would happen?",
    "2756020": "Nice find!  Thank you for sharing.",
    "2756142": "Looked at train_metadata.csv a little bit...\n\nThe couple duplicates with different species labels I looked at manually did -not- have secondary labels. \n\nsome of them at least had related species:\n\nwoosan/XC825766.ogg Wood Sandpiper\ngrnsan/XC825765.ogg Green Sandpiper\n\nwhbwoo2/XC239509.ogg White-bellied Woodpecker\nrufwoo2/XC239509.ogg Rufous Woodpecker\n\nSome of the dupes have different locations??? (lat / long)\n\"zitcis1/XC302781.ogg\" = 12.2714 \t109.1138 \t\n\"zitcis1/XC303866.ogg\" = 12.4419 \t109.2743\n\nSo - looking at the URLs - the listings show different sub-species?:\nhttps://xeno-canto.org/302781 Cisticola juncidis tinnabulans\nhttps://xeno-canto.org/303866  Cisticola juncidis\n\nI'll also note the recording quality on those last 2 was pretty bad (major background noise) - yet rated a 3?  Those 2 files clearly had the same audio - but were not byte-for-byte the same.\n\nI'm imagining someone correcting a misidentified species / moving stuff around (but leaving the original on accident)",
    "2756677": "Great investigation folks, thanks for the detective work! My bet would be that the (near) duplicates are partially remaining artifacts from taxonomy changes and (as you suspected) corrected species ID. Unfortunately, these are super hard to detect, and I believe they should be in the original dataset because this is the kind of data biologists have to deal with. There is no actual ground truth - sometimes there are inconsistencies. If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.\n\nIn terms of rating, these are also not super reliable, because they're based on user feedback (which is noisy at times).",
    "2759346": "> If you find a good and reliable way to identify near duplicates in audio - let us know, this does have some real-world application outside Kaggle.\n\nThresholding based on volume & embedding distance appears to be fairly robust, but the thresholds have to be tuned whenever changing the embedding model, and human curation afterwards is still required to find occasional false positives. Also I'm not sure how many other duplicates there are that this technique is not detecting.\n\nI'd be interested in exploring this further. Let me know if you have any ideas on directions to try :)",
    "2766563": "Nice job! Thank you!"
  },
  "source": "meta"
}