{
  "id": 396506,
  "title": "Duplicate audio files in the training data",
  "url": "/competitions/birdclef-2023/discussion/396506",
  "author_name": "",
  "post_date": "2023-03-21T21:01:39.766454Z",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Usually one of the first steps I take when exploring tabular data is to check for duplicates. I discovered that when the <code>filename</code> and <code>url</code> columns were dropped from <code>train_meta.csv</code> there were 2503 duplicate rows. No big deal right? Some of the files could have been recorded by the same <code>author</code>, in the same <code>location</code>, and have the same <code>rating</code> etc and still be totally different recordings. That's when I decided to dig a little deeper:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10590800%2Fd5a8b3f1f76a0c3aa612dedaf13c603a%2Fmel-spectrogram.png?generation=1679430426479307&amp;alt=media\" alt=\"\"></p>\n<p>I ended up finding 5 duplicate audio files in total, but there could definitely be more as I only explored the prementioned 2503 duplicate rows of <code>train_meta.csv</code>. Additionally, it appears that some of the audio files seem to be split into multiple parts which could potentially mean some of the audio files overlap each other. NOTE:  You can view all 5 duplicate pairs in the <a href=\"https://www.kaggle.com/code/mattop/birdclef-2023-eda/notebook\" target=\"_blank\">EDA notebook</a> I recently posted.</p>",
  "messages": [
    {
      "id": "2191282",
      "postDate": "03/21/2023 21:01:39",
      "content": "<p>Usually one of the first steps I take when exploring tabular data is to check for duplicates. I discovered that when the <code>filename</code> and <code>url</code> columns were dropped from <code>train_meta.csv</code> there were 2503 duplicate rows. No big deal right? Some of the files could have been recorded by the same <code>author</code>, in the same <code>location</code>, and have the same <code>rating</code> etc and still be totally different recordings. That's when I decided to dig a little deeper:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10590800%2Fd5a8b3f1f76a0c3aa612dedaf13c603a%2Fmel-spectrogram.png?generation=1679430426479307&amp;alt=media\" alt=\"\"></p>\n<p>I ended up finding 5 duplicate audio files in total, but there could definitely be more as I only explored the prementioned 2503 duplicate rows of <code>train_meta.csv</code>. Additionally, it appears that some of the audio files seem to be split into multiple parts which could potentially mean some of the audio files overlap each other. NOTE:  You can view all 5 duplicate pairs in the <a href=\"https://www.kaggle.com/code/mattop/birdclef-2023-eda/notebook\" target=\"_blank\">EDA notebook</a> I recently posted.</p>",
      "rawMarkdown": "Usually one of the first steps I take when exploring tabular data is to check for duplicates. I discovered that when the `filename` and `url` columns were dropped from `train_meta.csv` there were 2503 duplicate rows. No big deal right? Some of the files could have been recorded by the same `author`, in the same `location`, and have the same `rating` etc and still be totally different recordings. That's when I decided to dig a little deeper:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10590800%2Fd5a8b3f1f76a0c3aa612dedaf13c603a%2Fmel-spectrogram.png?generation=1679430426479307&alt=media)\n\nI ended up finding 5 duplicate audio files in total, but there could definitely be more as I only explored the prementioned 2503 duplicate rows of `train_meta.csv`. Additionally, it appears that some of the audio files seem to be split into multiple parts which could potentially mean some of the audio files overlap each other. NOTE:  You can view all 5 duplicate pairs in the [EDA notebook](https://www.kaggle.com/code/mattop/birdclef-2023-eda/notebook) I recently posted.",
      "votes": null
    },
    {
      "id": "2191830",
      "postDate": "03/22/2023 08:21:34",
      "content": "<p>Thanks for investigating. Yet, it seems there's not much we can do if duplicate files have a different XC ID (i.e., filename). Sometimes recordists might confuse files they upload, or (happens pretty often) might split a longer file into two before uploading.</p>",
      "rawMarkdown": "Thanks for investigating. Yet, it seems there's not much we can do if duplicate files have a different XC ID (i.e., filename). Sometimes recordists might confuse files they upload, or (happens pretty often) might split a longer file into two before uploading.",
      "votes": null
    },
    {
      "id": "2194363",
      "postDate": "03/23/2023 22:05:36",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, thank you for your response. My original strategy for identifying these duplicate audio files was rather inefficient and I have recently came up with a much more efficient method. The key is to compute the audio duration for each file in <code>train_meta.csv</code>. This can be achieved using <code>librosa.get_duration</code> with good floating point precision. When the audio length, <code>author</code> and location (<code>longitude</code>, <code>latitude</code>) are the same, there is potential for a duplicate audio file. Using this method I have discovered an additional 5 duplicate audio files (for a total of 10) and I have added them to my <a href=\"https://www.kaggle.com/code/mattop/birdclef-2023-eda\" target=\"_blank\">EDA notebook</a>. I have also created a dataset that is a copy of <code>train_meta.csv</code> but I added two additional columns: <code>duration_seconds</code> &amp; <code>duration_minutes</code> that you can find <a href=\"https://www.kaggle.com/datasets/mattop/birdclef-2023-train-meta-w-audio-durations\" target=\"_blank\">here</a> if anyone would like to attempt to find more duplicate audio files. </p>",
      "rawMarkdown": "Hello @stefankahl, thank you for your response. My original strategy for identifying these duplicate audio files was rather inefficient and I have recently came up with a much more efficient method. The key is to compute the audio duration for each file in `train_meta.csv`. This can be achieved using `librosa.get_duration` with good floating point precision. When the audio length, `author` and location (`longitude`, `latitude`) are the same, there is potential for a duplicate audio file. Using this method I have discovered an additional 5 duplicate audio files (for a total of 10) and I have added them to my [EDA notebook](https://www.kaggle.com/code/mattop/birdclef-2023-eda). I have also created a dataset that is a copy of `train_meta.csv` but I added two additional columns: `duration_seconds` & `duration_minutes` that you can find [here](https://www.kaggle.com/datasets/mattop/birdclef-2023-train-meta-w-audio-durations) if anyone would like to attempt to find more duplicate audio files.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2191830,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "03/22/2023 08:21:34",
      "content": "<p>Thanks for investigating. Yet, it seems there's not much we can do if duplicate files have a different XC ID (i.e., filename). Sometimes recordists might confuse files they upload, or (happens pretty often) might split a longer file into two before uploading.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2194363,
          "author_name": "mattop",
          "author_url": "",
          "post_date": "03/23/2023 22:05:36",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>, thank you for your response. My original strategy for identifying these duplicate audio files was rather inefficient and I have recently came up with a much more efficient method. The key is to compute the audio duration for each file in <code>train_meta.csv</code>. This can be achieved using <code>librosa.get_duration</code> with good floating point precision. When the audio length, <code>author</code> and location (<code>longitude</code>, <code>latitude</code>) are the same, there is potential for a duplicate audio file. Using this method I have discovered an additional 5 duplicate audio files (for a total of 10) and I have added them to my <a href=\"https://www.kaggle.com/code/mattop/birdclef-2023-eda\" target=\"_blank\">EDA notebook</a>. I have also created a dataset that is a copy of <code>train_meta.csv</code> but I added two additional columns: <code>duration_seconds</code> &amp; <code>duration_minutes</code> that you can find <a href=\"https://www.kaggle.com/datasets/mattop/birdclef-2023-train-meta-w-audio-durations\" target=\"_blank\">here</a> if anyone would like to attempt to find more duplicate audio files. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2191282": "Usually one of the first steps I take when exploring tabular data is to check for duplicates. I discovered that when the `filename` and `url` columns were dropped from `train_meta.csv` there were 2503 duplicate rows. No big deal right? Some of the files could have been recorded by the same `author`, in the same `location`, and have the same `rating` etc and still be totally different recordings. That's when I decided to dig a little deeper:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10590800%2Fd5a8b3f1f76a0c3aa612dedaf13c603a%2Fmel-spectrogram.png?generation=1679430426479307&alt=media)\n\nI ended up finding 5 duplicate audio files in total, but there could definitely be more as I only explored the prementioned 2503 duplicate rows of `train_meta.csv`. Additionally, it appears that some of the audio files seem to be split into multiple parts which could potentially mean some of the audio files overlap each other. NOTE:  You can view all 5 duplicate pairs in the [EDA notebook](https://www.kaggle.com/code/mattop/birdclef-2023-eda/notebook) I recently posted.",
    "2191830": "Thanks for investigating. Yet, it seems there's not much we can do if duplicate files have a different XC ID (i.e., filename). Sometimes recordists might confuse files they upload, or (happens pretty often) might split a longer file into two before uploading.",
    "2194363": "Hello @stefankahl, thank you for your response. My original strategy for identifying these duplicate audio files was rather inefficient and I have recently came up with a much more efficient method. The key is to compute the audio duration for each file in `train_meta.csv`. This can be achieved using `librosa.get_duration` with good floating point precision. When the audio length, `author` and location (`longitude`, `latitude`) are the same, there is potential for a duplicate audio file. Using this method I have discovered an additional 5 duplicate audio files (for a total of 10) and I have added them to my [EDA notebook](https://www.kaggle.com/code/mattop/birdclef-2023-eda). I have also created a dataset that is a copy of `train_meta.csv` but I added two additional columns: `duration_seconds` & `duration_minutes` that you can find [here](https://www.kaggle.com/datasets/mattop/birdclef-2023-train-meta-w-audio-durations) if anyone would like to attempt to find more duplicate audio files."
  },
  "source": "meta"
}