{
  "id": 161049,
  "title": "here's how to handle variable lengths",
  "url": "/competitions/birdsong-recognition/discussion/161049",
  "author_name": "",
  "post_date": "2020-06-23T15:19:45.732195300Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>If you are using images then they can be resized but what if you want to use audio features and some form of recurrent model then you need to handle different lengths of features. </p>\n\n<p>In NLP side we generally zero pad but given that this is sound classification problem we can do better.</p>\n\n<p>We can wrap the sound features as if all clips are played on repeat simultaneously and stoped at common timestep.</p>\n\n<p>This can be done efficiently using collect_fn function for pytorch dataloader and we do this at batch level instead of fixed length at dataset level cool right ?</p>\n\n<p>if you need code <a href=\"https://www.kaggle.com/dhananjay3/simple-pytorch-starter\">checkout my super simple starter notebook</a> which has just enough code to get you started with audio classification. </p>",
  "messages": [
    {
      "id": "898539",
      "postDate": "06/23/2020 15:19:45",
      "content": "<p>If you are using images then they can be resized but what if you want to use audio features and some form of recurrent model then you need to handle different lengths of features. </p>\n\n<p>In NLP side we generally zero pad but given that this is sound classification problem we can do better.</p>\n\n<p>We can wrap the sound features as if all clips are played on repeat simultaneously and stoped at common timestep.</p>\n\n<p>This can be done efficiently using collect_fn function for pytorch dataloader and we do this at batch level instead of fixed length at dataset level cool right ?</p>\n\n<p>if you need code <a href=\"https://www.kaggle.com/dhananjay3/simple-pytorch-starter\">checkout my super simple starter notebook</a> which has just enough code to get you started with audio classification. </p>",
      "rawMarkdown": "If you are using images then they can be resized but what if you want to use audio features and some form of recurrent model then you need to handle different lengths of features. \n\nIn NLP side we generally zero pad but given that this is sound classification problem we can do better.\n\n\nWe can wrap the sound features as if all clips are played on repeat simultaneously and stoped at common timestep.\n\n This can be done efficiently using collect_fn function for pytorch dataloader and we do this at batch level instead of fixed length at dataset level cool right ?\n\n\nif you need code [checkout my super simple starter notebook](https://www.kaggle.com/dhananjay3/simple-pytorch-starter) which has just enough code to get you started with audio classification.",
      "votes": null
    },
    {
      "id": "898845",
      "postDate": "06/23/2020 19:37:55",
      "content": "<p>If your image CNN model uses <code>GlobalAveragePooling2D()</code> as the last layer, then you don't need to repeat (tile) images. You can just feed in different sized images and it will work the same way as if you repeated the image.</p>",
      "rawMarkdown": "If your image CNN model uses `GlobalAveragePooling2D()` as the last layer, then you don't need to repeat (tile) images. You can just feed in different sized images and it will work the same way as if you repeated the image.",
      "votes": null
    },
    {
      "id": "899098",
      "postDate": "06/24/2020 01:38:50",
      "content": "<p><code>GlobalAveragePooling2D()</code> is used to handle variable dims across batches. but to create batches we need to resize features (when using images) to same dimensions which is unintuitive for different length clips.  </p>",
      "rawMarkdown": "`GlobalAveragePooling2D()` is used to handle variable dims across batches. but to create batches we need to resize features (when using images) to same dimensions which is unintuitive for different length clips.",
      "votes": null
    },
    {
      "id": "903425",
      "postDate": "06/26/2020 20:24:46",
      "content": "<p>Great information in that post mate! Kudos for that input! Keep sharing! Cheers 👍 </p>\n\n<p>Do check out <a href=\"https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe\">my notebook</a> for the competition on feature extraction. It involves a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you. </p>",
      "rawMarkdown": "Great information in that post mate! Kudos for that input! Keep sharing! Cheers 👍 \n\nDo check out [my notebook](https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe) for the competition on feature extraction. It involves a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 898845,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/23/2020 19:37:55",
      "content": "<p>If your image CNN model uses <code>GlobalAveragePooling2D()</code> as the last layer, then you don't need to repeat (tile) images. You can just feed in different sized images and it will work the same way as if you repeated the image.</p>",
      "votes": null,
      "replies": [
        {
          "id": 899098,
          "author_name": "dhananjay3",
          "author_url": "",
          "post_date": "06/24/2020 01:38:50",
          "content": "<p><code>GlobalAveragePooling2D()</code> is used to handle variable dims across batches. but to create batches we need to resize features (when using images) to same dimensions which is unintuitive for different length clips.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903425,
      "author_name": "navinmundhra",
      "author_url": "",
      "post_date": "06/26/2020 20:24:46",
      "content": "<p>Great information in that post mate! Kudos for that input! Keep sharing! Cheers 👍 </p>\n\n<p>Do check out <a href=\"https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe\">my notebook</a> for the competition on feature extraction. It involves a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "898539": "If you are using images then they can be resized but what if you want to use audio features and some form of recurrent model then you need to handle different lengths of features. \n\nIn NLP side we generally zero pad but given that this is sound classification problem we can do better.\n\n\nWe can wrap the sound features as if all clips are played on repeat simultaneously and stoped at common timestep.\n\n This can be done efficiently using collect_fn function for pytorch dataloader and we do this at batch level instead of fixed length at dataset level cool right ?\n\n\nif you need code [checkout my super simple starter notebook](https://www.kaggle.com/dhananjay3/simple-pytorch-starter) which has just enough code to get you started with audio classification.",
    "898845": "If your image CNN model uses `GlobalAveragePooling2D()` as the last layer, then you don't need to repeat (tile) images. You can just feed in different sized images and it will work the same way as if you repeated the image.",
    "899098": "`GlobalAveragePooling2D()` is used to handle variable dims across batches. but to create batches we need to resize features (when using images) to same dimensions which is unintuitive for different length clips.",
    "903425": "Great information in that post mate! Kudos for that input! Keep sharing! Cheers 👍 \n\nDo check out [my notebook](https://www.kaggle.com/navinmundhra/cornell-birdcall-extensive-eda-fe) for the competition on feature extraction. It involves a lot of visualizations, analysis, and hard work I put in for days to make sure I am providing sound and fundamental concepts. A bit of support would be appreciated if you enjoy it or you could let me know in the comments if you have some suggestions or criticism for me! :) Thank you."
  },
  "source": "meta"
}