{
  "id": 164364,
  "title": "Ways to deal with different audio lengths",
  "url": "/competitions/birdsong-recognition/discussion/164364",
  "author_name": "",
  "post_date": "2020-07-06T01:57:45.291689800Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I start this thread so people can share their ways to deal with the problem of audios having different lengths. My first approach has been using MFCCs. So far, I have been trying to pad the MFCCs to deal with different audio lengths. MFCCs are <strong>n_mfcc x frames</strong> (where n_mfcc is constant and the frames depend on the audio length), but since some audios have &gt;15000 frames the processed data is huge.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3197853%2F52f8ac05c02c1d7158317e15a954ac69%2FScreen%20Shot%202020-07-05%20at%2010.55.38%20PM.png?generation=1594000574152391&amp;alt=media\" alt=\"\"></p>\n\n<p>What is your way of dealing with this issue? </p>",
  "messages": [
    {
      "id": "916748",
      "postDate": "07/06/2020 01:57:45",
      "content": "<p>I start this thread so people can share their ways to deal with the problem of audios having different lengths. My first approach has been using MFCCs. So far, I have been trying to pad the MFCCs to deal with different audio lengths. MFCCs are <strong>n_mfcc x frames</strong> (where n_mfcc is constant and the frames depend on the audio length), but since some audios have &gt;15000 frames the processed data is huge.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3197853%2F52f8ac05c02c1d7158317e15a954ac69%2FScreen%20Shot%202020-07-05%20at%2010.55.38%20PM.png?generation=1594000574152391&amp;alt=media\" alt=\"\"></p>\n\n<p>What is your way of dealing with this issue? </p>",
      "rawMarkdown": "I start this thread so people can share their ways to deal with the problem of audios having different lengths. My first approach has been using MFCCs. So far, I have been trying to pad the MFCCs to deal with different audio lengths. MFCCs are **n_mfcc x frames** (where n_mfcc is constant and the frames depend on the audio length), but since some audios have &gt;15000 frames the processed data is huge.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3197853%2F52f8ac05c02c1d7158317e15a954ac69%2FScreen%20Shot%202020-07-05%20at%2010.55.38%20PM.png?generation=1594000574152391&amp;alt=media)\n\n\nWhat is your way of dealing with this issue?",
      "votes": null
    },
    {
      "id": "916766",
      "postDate": "07/06/2020 02:32:39",
      "content": "<p>Use a fully convolutional network and change your batch size to one. 👀 Alternatively, you could try resizing and trimming your training data.</p>",
      "rawMarkdown": "Use a fully convolutional network and change your batch size to one. 👀 Alternatively, you could try resizing and trimming your training data.",
      "votes": null
    },
    {
      "id": "916813",
      "postDate": "07/06/2020 04:05:43",
      "content": "<p>Padding and resizing mostly.</p>",
      "rawMarkdown": "Padding and resizing mostly.",
      "votes": null
    },
    {
      "id": "916997",
      "postDate": "07/06/2020 07:00:11",
      "content": "<p>You could split each audio file into 5-second chunks and compute the features for each segment. This way, you can use fixed input sizes for spectrograms or other features like MFCCs. The only thing you'd have to do is automatically assess if there's a signal (bird) in a chunk or just noise (static, wind, etc..) before using it for training.</p>",
      "rawMarkdown": "You could split each audio file into 5-second chunks and compute the features for each segment. This way, you can use fixed input sizes for spectrograms or other features like MFCCs. The only thing you'd have to do is automatically assess if there's a signal (bird) in a chunk or just noise (static, wind, etc..) before using it for training.",
      "votes": null
    },
    {
      "id": "917149",
      "postDate": "07/06/2020 09:15:46",
      "content": "<p>My approach here : <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/161049\">https://www.kaggle.com/c/birdsong-recognition/discussion/161049</a></p>",
      "rawMarkdown": "My approach here : https://www.kaggle.com/c/birdsong-recognition/discussion/161049",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 916766,
      "author_name": "alansun17904",
      "author_url": "",
      "post_date": "07/06/2020 02:32:39",
      "content": "<p>Use a fully convolutional network and change your batch size to one. 👀 Alternatively, you could try resizing and trimming your training data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 916813,
      "author_name": "jonykarki",
      "author_url": "",
      "post_date": "07/06/2020 04:05:43",
      "content": "<p>Padding and resizing mostly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 916997,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "07/06/2020 07:00:11",
      "content": "<p>You could split each audio file into 5-second chunks and compute the features for each segment. This way, you can use fixed input sizes for spectrograms or other features like MFCCs. The only thing you'd have to do is automatically assess if there's a signal (bird) in a chunk or just noise (static, wind, etc..) before using it for training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 917149,
      "author_name": "dhananjay3",
      "author_url": "",
      "post_date": "07/06/2020 09:15:46",
      "content": "<p>My approach here : <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/161049\">https://www.kaggle.com/c/birdsong-recognition/discussion/161049</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "916748": "I start this thread so people can share their ways to deal with the problem of audios having different lengths. My first approach has been using MFCCs. So far, I have been trying to pad the MFCCs to deal with different audio lengths. MFCCs are **n_mfcc x frames** (where n_mfcc is constant and the frames depend on the audio length), but since some audios have &gt;15000 frames the processed data is huge.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3197853%2F52f8ac05c02c1d7158317e15a954ac69%2FScreen%20Shot%202020-07-05%20at%2010.55.38%20PM.png?generation=1594000574152391&amp;alt=media)\n\n\nWhat is your way of dealing with this issue?",
    "916766": "Use a fully convolutional network and change your batch size to one. 👀 Alternatively, you could try resizing and trimming your training data.",
    "916813": "Padding and resizing mostly.",
    "916997": "You could split each audio file into 5-second chunks and compute the features for each segment. This way, you can use fixed input sizes for spectrograms or other features like MFCCs. The only thing you'd have to do is automatically assess if there's a signal (bird) in a chunk or just noise (static, wind, etc..) before using it for training.",
    "917149": "My approach here : https://www.kaggle.com/c/birdsong-recognition/discussion/161049"
  },
  "source": "meta"
}