{
  "id": 160260,
  "title": "how to best represent a vocalization?",
  "url": "/competitions/birdsong-recognition/discussion/160260",
  "author_name": "",
  "post_date": "2020-06-20T14:14:50.434261900Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm sure soon enough we will start training on spectrograms 🙂 </p>\n\n<p>The question is - how do we best represent the vocalization for classification? Is 5 second window optimal? Would we get better results if for example we instead looked at 5 x 1 second slices instead?</p>\n\n<p><a href=\"https://github.com/DrCoffey/DeepSqueak\">DeepSqueak</a> does something really cool. The create what they refer to as <code>sonograms</code> but my understanding is that this just a spectrogram with some interesting preprocessing applied</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fb38e86f9ea546bc463e13c9764dbb34b%2Fsonogram.png?generation=1592661752677119&amp;alt=media\" alt=\"\"></p>\n\n<p>I would really, really like to put my hands on a representation such as this. Thinking this might be (one) of the keys to this competition. The code is in matlab but the magic is also explained in the paper <a href=\"https://www.nature.com/articles/s41386-018-0303-6\">here</a>.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F19107001775e73114fc8e6e03977a49e%2Fcontour_detection.png?generation=1592661928101935&amp;alt=media\" alt=\"\"></p>\n\n<p>There is also some additional work on this in this <a href=\"https://github.com/mvansegbroeck/mupet\">repo</a> with a gammatone representation.</p>\n\n<p>I will update the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160222\">starter pack</a> soon-ish to run on spectrograms. But if we have any math wizzes among our mids, especially should they be familiar with matlab  and scipy or matplotlib on Python's side, I think figuring out how to create such sonograms would be super, super valuable.</p>\n\n<p>My most esteemed colleague from work is working on an <a href=\"https://github.com/earthspecies/representation-toolbox\">acoustic representation toolbox</a>, putting together a bunch of ways to represent sound, so I am thinking this can also be very helpful for the competition. Some really exotic transforms there so worth checking out 😉</p>\n\n<p>Anyhow - this competition is so much fun! 😊 Any leads, thoughts, ideas or code on those sonograms from deepsqueak or mupet would be extremely welcome!!!!!!!</p>\n\n<p>BTW I understand we expect a lot of ambient sounds in test, but I am thinking that maybe, just maybe, this sonogram technique could make the vocalizations stand out nonetheless.</p>",
  "messages": [
    {
      "id": "894548",
      "postDate": "06/20/2020 14:14:50",
      "content": "<p>I'm sure soon enough we will start training on spectrograms 🙂 </p>\n\n<p>The question is - how do we best represent the vocalization for classification? Is 5 second window optimal? Would we get better results if for example we instead looked at 5 x 1 second slices instead?</p>\n\n<p><a href=\"https://github.com/DrCoffey/DeepSqueak\">DeepSqueak</a> does something really cool. The create what they refer to as <code>sonograms</code> but my understanding is that this just a spectrogram with some interesting preprocessing applied</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fb38e86f9ea546bc463e13c9764dbb34b%2Fsonogram.png?generation=1592661752677119&amp;alt=media\" alt=\"\"></p>\n\n<p>I would really, really like to put my hands on a representation such as this. Thinking this might be (one) of the keys to this competition. The code is in matlab but the magic is also explained in the paper <a href=\"https://www.nature.com/articles/s41386-018-0303-6\">here</a>.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F19107001775e73114fc8e6e03977a49e%2Fcontour_detection.png?generation=1592661928101935&amp;alt=media\" alt=\"\"></p>\n\n<p>There is also some additional work on this in this <a href=\"https://github.com/mvansegbroeck/mupet\">repo</a> with a gammatone representation.</p>\n\n<p>I will update the <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160222\">starter pack</a> soon-ish to run on spectrograms. But if we have any math wizzes among our mids, especially should they be familiar with matlab  and scipy or matplotlib on Python's side, I think figuring out how to create such sonograms would be super, super valuable.</p>\n\n<p>My most esteemed colleague from work is working on an <a href=\"https://github.com/earthspecies/representation-toolbox\">acoustic representation toolbox</a>, putting together a bunch of ways to represent sound, so I am thinking this can also be very helpful for the competition. Some really exotic transforms there so worth checking out 😉</p>\n\n<p>Anyhow - this competition is so much fun! 😊 Any leads, thoughts, ideas or code on those sonograms from deepsqueak or mupet would be extremely welcome!!!!!!!</p>\n\n<p>BTW I understand we expect a lot of ambient sounds in test, but I am thinking that maybe, just maybe, this sonogram technique could make the vocalizations stand out nonetheless.</p>",
      "rawMarkdown": "I'm sure soon enough we will start training on spectrograms 🙂 \n\nThe question is - how do we best represent the vocalization for classification? Is 5 second window optimal? Would we get better results if for example we instead looked at 5 x 1 second slices instead?\n\n[DeepSqueak](https://github.com/DrCoffey/DeepSqueak) does something really cool. The create what they refer to as `sonograms` but my understanding is that this just a spectrogram with some interesting preprocessing applied\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fb38e86f9ea546bc463e13c9764dbb34b%2Fsonogram.png?generation=1592661752677119&amp;alt=media)\n\nI would really, really like to put my hands on a representation such as this. Thinking this might be (one) of the keys to this competition. The code is in matlab but the magic is also explained in the paper [here](https://www.nature.com/articles/s41386-018-0303-6).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F19107001775e73114fc8e6e03977a49e%2Fcontour_detection.png?generation=1592661928101935&amp;alt=media)\n\nThere is also some additional work on this in this [repo](https://github.com/mvansegbroeck/mupet) with a gammatone representation.\n\nI will update the [starter pack](https://www.kaggle.com/c/birdsong-recognition/discussion/160222) soon-ish to run on spectrograms. But if we have any math wizzes among our mids, especially should they be familiar with matlab  and scipy or matplotlib on Python's side, I think figuring out how to create such sonograms would be super, super valuable.\n\nMy most esteemed colleague from work is working on an [acoustic representation toolbox](https://github.com/earthspecies/representation-toolbox), putting together a bunch of ways to represent sound, so I am thinking this can also be very helpful for the competition. Some really exotic transforms there so worth checking out 😉\n\nAnyhow - this competition is so much fun! 😊 Any leads, thoughts, ideas or code on those sonograms from deepsqueak or mupet would be extremely welcome!!!!!!!\n\nBTW I understand we expect a lot of ambient sounds in test, but I am thinking that maybe, just maybe, this sonogram technique could make the vocalizations stand out nonetheless.",
      "votes": null
    },
    {
      "id": "894811",
      "postDate": "06/20/2020 19:16:05",
      "content": "<p>Thanks for sharing your ideas <a href=\"/radek1\">@radek1</a> :) Spectrograms worked very well in previous competitions like freesound-audio-tagging-2019. Here the audio samples are much longer I'm still thinking what should be the best way to work with such spectrograms. For now I'm just random sampling portions of each spectrogram. </p>\n\n<p>I'm using this function as item_tfms in fastai 2 datablock:</p>\n\n<p><code>\ndef time_slice(x:PILImage, size=128):\n    width, height = x.size[-2:]\n    start = np.random.randint(0, width-size) if width-size &gt; 0 else 0\n    x = x.crop((start,0,start+size,height))\n    return x\n</code></p>\n\n<p>To get square images like this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Ffbe4e128340e4e5820feff8b14c00039%2F__results___7_0.png?generation=1592680448245775&amp;alt=media\" alt=\"\"></p>\n\n<p>But the slice may not have any birdcall or even the incorrect one. Maybe looking for wider images can be better in this problem?</p>",
      "rawMarkdown": "Thanks for sharing your ideas @radek1 :) Spectrograms worked very well in previous competitions like freesound-audio-tagging-2019. Here the audio samples are much longer I'm still thinking what should be the best way to work with such spectrograms. For now I'm just random sampling portions of each spectrogram. \n\nI'm using this function as item_tfms in fastai 2 datablock:\n\n```\ndef time_slice(x:PILImage, size=128):\n    width, height = x.size[-2:]\n    start = np.random.randint(0, width-size) if width-size &gt; 0 else 0\n    x = x.crop((start,0,start+size,height))\n    return x\n```\n\nTo get square images like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Ffbe4e128340e4e5820feff8b14c00039%2F__results___7_0.png?generation=1592680448245775&amp;alt=media)\n\nBut the slice may not have any birdcall or even the incorrect one. Maybe looking for wider images can be better in this problem?",
      "votes": null
    },
    {
      "id": "894832",
      "postDate": "06/20/2020 19:35:21",
      "content": "<p>Agree, there is definitely not a clear cut answer how to approach this 🙂 One thing I was thinking about was grabbing 5 continuous seconds and slicing them up into 1 sec windows and creating a spectrogram on each, then feeding this into the model as 5 separate channels. My reasoning is that it might be tough to squish a 5 second spectrogram into a square without losing too much information. Alternatively one could try training on wide examples, 64x320 for instance or something like that, but that might also come with a set of challenges.</p>\n\n<p>Also, this is just my intuition and I might be wrong on that one, but for a deep learning model I think it shouldn't be too big of a problem if we sample data during train and every now and then miss the vocalization. We could also try label smoothing as an attempt to address this.</p>\n\n<p>Anyhow, please take these with a grain of salt, just a bunch of thoughts 🙂 Will be interesting to see what results people get with various approaches.</p>",
      "rawMarkdown": "Agree, there is definitely not a clear cut answer how to approach this 🙂 One thing I was thinking about was grabbing 5 continuous seconds and slicing them up into 1 sec windows and creating a spectrogram on each, then feeding this into the model as 5 separate channels. My reasoning is that it might be tough to squish a 5 second spectrogram into a square without losing too much information. Alternatively one could try training on wide examples, 64x320 for instance or something like that, but that might also come with a set of challenges.\n\nAlso, this is just my intuition and I might be wrong on that one, but for a deep learning model I think it shouldn't be too big of a problem if we sample data during train and every now and then miss the vocalization. We could also try label smoothing as an attempt to address this.\n\nAnyhow, please take these with a grain of salt, just a bunch of thoughts 🙂 Will be interesting to see what results people get with various approaches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 894811,
      "author_name": "mnpinto",
      "author_url": "",
      "post_date": "06/20/2020 19:16:05",
      "content": "<p>Thanks for sharing your ideas <a href=\"/radek1\">@radek1</a> :) Spectrograms worked very well in previous competitions like freesound-audio-tagging-2019. Here the audio samples are much longer I'm still thinking what should be the best way to work with such spectrograms. For now I'm just random sampling portions of each spectrogram. </p>\n\n<p>I'm using this function as item_tfms in fastai 2 datablock:</p>\n\n<p><code>\ndef time_slice(x:PILImage, size=128):\n    width, height = x.size[-2:]\n    start = np.random.randint(0, width-size) if width-size &gt; 0 else 0\n    x = x.crop((start,0,start+size,height))\n    return x\n</code></p>\n\n<p>To get square images like this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Ffbe4e128340e4e5820feff8b14c00039%2F__results___7_0.png?generation=1592680448245775&amp;alt=media\" alt=\"\"></p>\n\n<p>But the slice may not have any birdcall or even the incorrect one. Maybe looking for wider images can be better in this problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 894832,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "06/20/2020 19:35:21",
          "content": "<p>Agree, there is definitely not a clear cut answer how to approach this 🙂 One thing I was thinking about was grabbing 5 continuous seconds and slicing them up into 1 sec windows and creating a spectrogram on each, then feeding this into the model as 5 separate channels. My reasoning is that it might be tough to squish a 5 second spectrogram into a square without losing too much information. Alternatively one could try training on wide examples, 64x320 for instance or something like that, but that might also come with a set of challenges.</p>\n\n<p>Also, this is just my intuition and I might be wrong on that one, but for a deep learning model I think it shouldn't be too big of a problem if we sample data during train and every now and then miss the vocalization. We could also try label smoothing as an attempt to address this.</p>\n\n<p>Anyhow, please take these with a grain of salt, just a bunch of thoughts 🙂 Will be interesting to see what results people get with various approaches.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "894548": "I'm sure soon enough we will start training on spectrograms 🙂 \n\nThe question is - how do we best represent the vocalization for classification? Is 5 second window optimal? Would we get better results if for example we instead looked at 5 x 1 second slices instead?\n\n[DeepSqueak](https://github.com/DrCoffey/DeepSqueak) does something really cool. The create what they refer to as `sonograms` but my understanding is that this just a spectrogram with some interesting preprocessing applied\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2Fb38e86f9ea546bc463e13c9764dbb34b%2Fsonogram.png?generation=1592661752677119&amp;alt=media)\n\nI would really, really like to put my hands on a representation such as this. Thinking this might be (one) of the keys to this competition. The code is in matlab but the magic is also explained in the paper [here](https://www.nature.com/articles/s41386-018-0303-6).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F83267%2F19107001775e73114fc8e6e03977a49e%2Fcontour_detection.png?generation=1592661928101935&amp;alt=media)\n\nThere is also some additional work on this in this [repo](https://github.com/mvansegbroeck/mupet) with a gammatone representation.\n\nI will update the [starter pack](https://www.kaggle.com/c/birdsong-recognition/discussion/160222) soon-ish to run on spectrograms. But if we have any math wizzes among our mids, especially should they be familiar with matlab  and scipy or matplotlib on Python's side, I think figuring out how to create such sonograms would be super, super valuable.\n\nMy most esteemed colleague from work is working on an [acoustic representation toolbox](https://github.com/earthspecies/representation-toolbox), putting together a bunch of ways to represent sound, so I am thinking this can also be very helpful for the competition. Some really exotic transforms there so worth checking out 😉\n\nAnyhow - this competition is so much fun! 😊 Any leads, thoughts, ideas or code on those sonograms from deepsqueak or mupet would be extremely welcome!!!!!!!\n\nBTW I understand we expect a lot of ambient sounds in test, but I am thinking that maybe, just maybe, this sonogram technique could make the vocalizations stand out nonetheless.",
    "894811": "Thanks for sharing your ideas @radek1 :) Spectrograms worked very well in previous competitions like freesound-audio-tagging-2019. Here the audio samples are much longer I'm still thinking what should be the best way to work with such spectrograms. For now I'm just random sampling portions of each spectrogram. \n\nI'm using this function as item_tfms in fastai 2 datablock:\n\n```\ndef time_slice(x:PILImage, size=128):\n    width, height = x.size[-2:]\n    start = np.random.randint(0, width-size) if width-size &gt; 0 else 0\n    x = x.crop((start,0,start+size,height))\n    return x\n```\n\nTo get square images like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Ffbe4e128340e4e5820feff8b14c00039%2F__results___7_0.png?generation=1592680448245775&amp;alt=media)\n\nBut the slice may not have any birdcall or even the incorrect one. Maybe looking for wider images can be better in this problem?",
    "894832": "Agree, there is definitely not a clear cut answer how to approach this 🙂 One thing I was thinking about was grabbing 5 continuous seconds and slicing them up into 1 sec windows and creating a spectrogram on each, then feeding this into the model as 5 separate channels. My reasoning is that it might be tough to squish a 5 second spectrogram into a square without losing too much information. Alternatively one could try training on wide examples, 64x320 for instance or something like that, but that might also come with a set of challenges.\n\nAlso, this is just my intuition and I might be wrong on that one, but for a deep learning model I think it shouldn't be too big of a problem if we sample data during train and every now and then miss the vocalization. We could also try label smoothing as an attempt to address this.\n\nAnyhow, please take these with a grain of salt, just a bunch of thoughts 🙂 Will be interesting to see what results people get with various approaches."
  },
  "source": "meta"
}