{
  "id": 177547,
  "title": "Strange durations",
  "url": "/competitions/birdsong-recognition/discussion/177547",
  "author_name": "",
  "post_date": "2020-08-26T09:41:12.857813800Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am entering this competition hence I have not yet read all the forum.  If this has already been discussed then just let me know.</p>\n<p>I computed every clip duration by dividing actual clip length by sampling rate, and I found some weirdness.  In most cases the duration reported in train.csv is the actual duration rounded to closest lower integer, but there are some weird outliers as shown in this scatter plot.  x is actual duration, y is difference between reported duration and actual duration, in seconds.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fc32a83a87bfabb84f37d5bc28b6c4b85%2Fduration.png?generation=1598434548907532&amp;alt=media\" alt=\"\"></p>\n<p>The 6 outliers for which the difference is at least 2 seconds are:</p>\n<pre><code>     ebird_code  filename\n 4892     canwar  XC131174.mp3\n 9879     haiwoo  XC412239.mp3\n10224     hoowar  XC134259.mp3\n12211     lotduc  XC195038.mp3\n12262     louwat  XC134962.mp3\n18319     swathr  XC129847.mp3\n</code></pre>\n<p>Actual sampling rate is the same as the one reported in train.csv except for one file where the sampling rate in the file is 0.  It is this file:</p>\n<pre><code>    ebird_code  filename\n12211     lotduc  XC195038.mp3\n</code></pre>\n<p>I used librosa with ffmpeg to read mp3 files.</p>",
  "messages": [
    {
      "id": "986206",
      "postDate": "08/26/2020 09:41:12",
      "content": "<p>I am entering this competition hence I have not yet read all the forum.  If this has already been discussed then just let me know.</p>\n<p>I computed every clip duration by dividing actual clip length by sampling rate, and I found some weirdness.  In most cases the duration reported in train.csv is the actual duration rounded to closest lower integer, but there are some weird outliers as shown in this scatter plot.  x is actual duration, y is difference between reported duration and actual duration, in seconds.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fc32a83a87bfabb84f37d5bc28b6c4b85%2Fduration.png?generation=1598434548907532&amp;alt=media\" alt=\"\"></p>\n<p>The 6 outliers for which the difference is at least 2 seconds are:</p>\n<pre><code>     ebird_code  filename\n 4892     canwar  XC131174.mp3\n 9879     haiwoo  XC412239.mp3\n10224     hoowar  XC134259.mp3\n12211     lotduc  XC195038.mp3\n12262     louwat  XC134962.mp3\n18319     swathr  XC129847.mp3\n</code></pre>\n<p>Actual sampling rate is the same as the one reported in train.csv except for one file where the sampling rate in the file is 0.  It is this file:</p>\n<pre><code>    ebird_code  filename\n12211     lotduc  XC195038.mp3\n</code></pre>\n<p>I used librosa with ffmpeg to read mp3 files.</p>",
      "rawMarkdown": "I am entering this competition hence I have not yet read all the forum.  If this has already been discussed then just let me know.\n\nI computed every clip duration by dividing actual clip length by sampling rate, and I found some weirdness.  In most cases the duration reported in train.csv is the actual duration rounded to closest lower integer, but there are some weird outliers as shown in this scatter plot.  x is actual duration, y is difference between reported duration and actual duration, in seconds.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fc32a83a87bfabb84f37d5bc28b6c4b85%2Fduration.png?generation=1598434548907532&alt=media)\n\nThe 6 outliers for which the difference is at least 2 seconds are:\n\n```\n \tebird_code \tfilename\n 4892 \tcanwar \tXC131174.mp3\n 9879 \thaiwoo \tXC412239.mp3\n10224 \thoowar \tXC134259.mp3\n12211 \tlotduc \tXC195038.mp3\n12262 \tlouwat \tXC134962.mp3\n18319 \tswathr \tXC129847.mp3\n```\n\nActual sampling rate is the same as the one reported in train.csv except for one file where the sampling rate in the file is 0.  It is this file:\n\n ```\n\tebird_code \tfilename\n12211 \tlotduc \tXC195038.mp3\n```\n\nI used librosa with ffmpeg to read mp3 files.",
      "votes": null
    },
    {
      "id": "986432",
      "postDate": "08/26/2020 13:45:33",
      "content": "<p>Yes I had to exclude 0 sampling rate file For training. Could not understand though what exactly zero sampling rate means. </p>",
      "rawMarkdown": "Yes I had to exclude 0 sampling rate file For training. Could not understand though what exactly zero sampling rate means.",
      "votes": null
    },
    {
      "id": "986439",
      "postDate": "08/26/2020 13:52:20",
      "content": "<p>you can use the sampling rate from train.csv</p>",
      "rawMarkdown": "you can use the sampling rate from train.csv",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 986432,
      "author_name": "shikha130vv",
      "author_url": "",
      "post_date": "08/26/2020 13:45:33",
      "content": "<p>Yes I had to exclude 0 sampling rate file For training. Could not understand though what exactly zero sampling rate means. </p>",
      "votes": null,
      "replies": [
        {
          "id": 986439,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/26/2020 13:52:20",
          "content": "<p>you can use the sampling rate from train.csv</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "986206": "I am entering this competition hence I have not yet read all the forum.  If this has already been discussed then just let me know.\n\nI computed every clip duration by dividing actual clip length by sampling rate, and I found some weirdness.  In most cases the duration reported in train.csv is the actual duration rounded to closest lower integer, but there are some weird outliers as shown in this scatter plot.  x is actual duration, y is difference between reported duration and actual duration, in seconds.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2Fc32a83a87bfabb84f37d5bc28b6c4b85%2Fduration.png?generation=1598434548907532&alt=media)\n\nThe 6 outliers for which the difference is at least 2 seconds are:\n\n```\n \tebird_code \tfilename\n 4892 \tcanwar \tXC131174.mp3\n 9879 \thaiwoo \tXC412239.mp3\n10224 \thoowar \tXC134259.mp3\n12211 \tlotduc \tXC195038.mp3\n12262 \tlouwat \tXC134962.mp3\n18319 \tswathr \tXC129847.mp3\n```\n\nActual sampling rate is the same as the one reported in train.csv except for one file where the sampling rate in the file is 0.  It is this file:\n\n ```\n\tebird_code \tfilename\n12211 \tlotduc \tXC195038.mp3\n```\n\nI used librosa with ffmpeg to read mp3 files.",
    "986432": "Yes I had to exclude 0 sampling rate file For training. Could not understand though what exactly zero sampling rate means.",
    "986439": "you can use the sampling rate from train.csv"
  },
  "source": "meta"
}