{
  "id": 175657,
  "title": "Using a nearest neighbors approach with a filtered spectrogram. ",
  "url": "/competitions/birdsong-recognition/discussion/175657",
  "author_name": "",
  "post_date": "2020-08-18T23:58:59.492760200Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I've essentially recreated the old Shazam algorithm that creates a \"constellation\" spectrogram, essentially a sparse matrix. </p>\n<p>I want to try fingerprinting all of the training data and then use a nearest neighbors approach for a baseline model. </p>\n<p>I'm concerned about how to go about loading so much data at once to fit sklearn's <code>KNeighborsClassifier</code>. Is this too ambitious? Sorry for the newb question. I'm new to handling such large datasets, as I'm recently coming from the beginner competitions like MNIST/Titanic/HomePrices. </p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "976543",
      "postDate": "08/18/2020 23:58:59",
      "content": "<p>I've essentially recreated the old Shazam algorithm that creates a \"constellation\" spectrogram, essentially a sparse matrix. </p>\n<p>I want to try fingerprinting all of the training data and then use a nearest neighbors approach for a baseline model. </p>\n<p>I'm concerned about how to go about loading so much data at once to fit sklearn's <code>KNeighborsClassifier</code>. Is this too ambitious? Sorry for the newb question. I'm new to handling such large datasets, as I'm recently coming from the beginner competitions like MNIST/Titanic/HomePrices. </p>\n<p>Thanks!</p>",
      "rawMarkdown": "I've essentially recreated the old Shazam algorithm that creates a \"constellation\" spectrogram, essentially a sparse matrix. \n\nI want to try fingerprinting all of the training data and then use a nearest neighbors approach for a baseline model. \n\nI'm concerned about how to go about loading so much data at once to fit sklearn's `KNeighborsClassifier`. Is this too ambitious? Sorry for the newb question. I'm new to handling such large datasets, as I'm recently coming from the beginner competitions like MNIST/Titanic/HomePrices. \n\nThanks!",
      "votes": null
    },
    {
      "id": "976668",
      "postDate": "08/19/2020 02:56:16",
      "content": "<p>I believe that you can train KNeighbors with the kaggle gpu, which might be able to handle all of the data. </p>",
      "rawMarkdown": "I believe that you can train KNeighbors with the kaggle gpu, which might be able to handle all of the data.",
      "votes": null
    },
    {
      "id": "976859",
      "postDate": "08/19/2020 06:26:14",
      "content": "<p>I'm also a bit of a newbie - so disclaimer - the below may be wrong!</p>\n<p>Going through this, I've tried to get features out of a semi-trained Resnet model (so just fed in the data to the model without really trying to make it accurate). So the data 'out' of this should have been somewhat representative of how 'similar' the calls were, some of their usual characteristics. I think it's usually termed embeddings, but I could be wrong.</p>\n<p>Running TSNE on this was too tough - as each audio file has to have multiple entries for different time points - it wasn't really viable to fit everything into memory. Should probably have extracted less 'feature' data per row. Anyhow, I managed to get PCA to run (as at least you can fit on some of the data, then transform the rest). It did give some interesting starting points for further analysis, but it's worth knowing that although clips are labelled as 'bird X', in fact 'bird X' may have a whole variety of different calls, and there may be other birds singing as well. And I'm not even sure all the 'official' labels for every track are absolutely 100% correct - others who know more than me may have a view - it would probably not be surprising if there's a bit of a grey area after all this is very hard to identify correctly. But mostly I think the issue is other birds singing + variability of each bird type.</p>\n<p>Just mentioning this because maybe I went about it the wrong way but it seemed like a large and difficult task to extract some basic separating features across all clips. It did give me some useful info but was much messier than I'd anticipated.</p>\n<p>Edit - there is of course also just a lot of blank space, background noise in the audio files.</p>",
      "rawMarkdown": "I'm also a bit of a newbie - so disclaimer - the below may be wrong!\n\nGoing through this, I've tried to get features out of a semi-trained Resnet model (so just fed in the data to the model without really trying to make it accurate). So the data 'out' of this should have been somewhat representative of how 'similar' the calls were, some of their usual characteristics. I think it's usually termed embeddings, but I could be wrong.\n\nRunning TSNE on this was too tough - as each audio file has to have multiple entries for different time points - it wasn't really viable to fit everything into memory. Should probably have extracted less 'feature' data per row. Anyhow, I managed to get PCA to run (as at least you can fit on some of the data, then transform the rest). It did give some interesting starting points for further analysis, but it's worth knowing that although clips are labelled as 'bird X', in fact 'bird X' may have a whole variety of different calls, and there may be other birds singing as well. And I'm not even sure all the 'official' labels for every track are absolutely 100% correct - others who know more than me may have a view - it would probably not be surprising if there's a bit of a grey area after all this is very hard to identify correctly. But mostly I think the issue is other birds singing + variability of each bird type.\n\nJust mentioning this because maybe I went about it the wrong way but it seemed like a large and difficult task to extract some basic separating features across all clips. It did give me some useful info but was much messier than I'd anticipated.\n\nEdit - there is of course also just a lot of blank space, background noise in the audio files.",
      "votes": null
    },
    {
      "id": "977279",
      "postDate": "08/19/2020 11:43:34",
      "content": "<p>Take a look at RAPIDS! Chris recently did some clustering on 2000+ dimensional embeddings of a few 10000's of images. It only took a few seconds :)</p>",
      "rawMarkdown": "Take a look at RAPIDS! Chris recently did some clustering on 2000+ dimensional embeddings of a few 10000's of images. It only took a few seconds :)",
      "votes": null
    },
    {
      "id": "977708",
      "postDate": "08/19/2020 16:54:24",
      "content": "<p>Hey, I'll check it out, thanks! Someone else also recommended checking out Dask. Any thoughts on one vs the other (from the perspective of a relative beginner who joined the comp with less than a month left)?</p>",
      "rawMarkdown": "Hey, I'll check it out, thanks! Someone else also recommended checking out Dask. Any thoughts on one vs the other (from the perspective of a relative beginner who joined the comp with less than a month left)?",
      "votes": null
    },
    {
      "id": "977717",
      "postDate": "08/19/2020 17:00:51",
      "content": "<p>Dask is for distributed computing as far as I know while RAPIDS works entirely on GPU, so I think that's still orders of magnitude faster. Especially if you are using Kaggle Notebooks. </p>",
      "rawMarkdown": "Dask is for distributed computing as far as I know while RAPIDS works entirely on GPU, so I think that's still orders of magnitude faster. Especially if you are using Kaggle Notebooks.",
      "votes": null
    },
    {
      "id": "977784",
      "postDate": "08/19/2020 17:46:27",
      "content": "<p>Thank you, I appreciate the insight. Good luck to you! </p>",
      "rawMarkdown": "Thank you, I appreciate the insight. Good luck to you!",
      "votes": null
    },
    {
      "id": "978096",
      "postDate": "08/20/2020 00:00:54",
      "content": "<p>you may want to take a look at deep learning based clustering that could work quite well</p>\n<p>Deep Clustering for Unsupervised Learning of Visual Features<br>\n<a href=\"https://github.com/facebookresearch/deepcluster\" target=\"_blank\">https://github.com/facebookresearch/deepcluster</a></p>\n<p>Self-labelling via simultaneous clustering and representation learning<br>\n<a href=\"https://github.com/yukimasano/self-label\" target=\"_blank\">https://github.com/yukimasano/self-label</a></p>",
      "rawMarkdown": "you may want to take a look at deep learning based clustering that could work quite well\n\nDeep Clustering for Unsupervised Learning of Visual Features\nhttps://github.com/facebookresearch/deepcluster\n\nSelf-labelling via simultaneous clustering and representation learning\nhttps://github.com/yukimasano/self-label",
      "votes": null
    },
    {
      "id": "992029",
      "postDate": "08/30/2020 20:32:19",
      "content": "<p>And you can use RAPIDS with Dask!</p>",
      "rawMarkdown": "And you can use RAPIDS with Dask!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 976859,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "08/19/2020 06:26:14",
      "content": "<p>I'm also a bit of a newbie - so disclaimer - the below may be wrong!</p>\n<p>Going through this, I've tried to get features out of a semi-trained Resnet model (so just fed in the data to the model without really trying to make it accurate). So the data 'out' of this should have been somewhat representative of how 'similar' the calls were, some of their usual characteristics. I think it's usually termed embeddings, but I could be wrong.</p>\n<p>Running TSNE on this was too tough - as each audio file has to have multiple entries for different time points - it wasn't really viable to fit everything into memory. Should probably have extracted less 'feature' data per row. Anyhow, I managed to get PCA to run (as at least you can fit on some of the data, then transform the rest). It did give some interesting starting points for further analysis, but it's worth knowing that although clips are labelled as 'bird X', in fact 'bird X' may have a whole variety of different calls, and there may be other birds singing as well. And I'm not even sure all the 'official' labels for every track are absolutely 100% correct - others who know more than me may have a view - it would probably not be surprising if there's a bit of a grey area after all this is very hard to identify correctly. But mostly I think the issue is other birds singing + variability of each bird type.</p>\n<p>Just mentioning this because maybe I went about it the wrong way but it seemed like a large and difficult task to extract some basic separating features across all clips. It did give me some useful info but was much messier than I'd anticipated.</p>\n<p>Edit - there is of course also just a lot of blank space, background noise in the audio files.</p>",
      "votes": null,
      "replies": [
        {
          "id": 977784,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/19/2020 17:46:27",
          "content": "<p>Thank you, I appreciate the insight. Good luck to you! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 977279,
      "author_name": "group16",
      "author_url": "",
      "post_date": "08/19/2020 11:43:34",
      "content": "<p>Take a look at RAPIDS! Chris recently did some clustering on 2000+ dimensional embeddings of a few 10000's of images. It only took a few seconds :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 977708,
          "author_name": "wesleyneill",
          "author_url": "",
          "post_date": "08/19/2020 16:54:24",
          "content": "<p>Hey, I'll check it out, thanks! Someone else also recommended checking out Dask. Any thoughts on one vs the other (from the perspective of a relative beginner who joined the comp with less than a month left)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 977717,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/19/2020 17:00:51",
          "content": "<p>Dask is for distributed computing as far as I know while RAPIDS works entirely on GPU, so I think that's still orders of magnitude faster. Especially if you are using Kaggle Notebooks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 992029,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/30/2020 20:32:19",
          "content": "<p>And you can use RAPIDS with Dask!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 978096,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/20/2020 00:00:54",
      "content": "<p>you may want to take a look at deep learning based clustering that could work quite well</p>\n<p>Deep Clustering for Unsupervised Learning of Visual Features<br>\n<a href=\"https://github.com/facebookresearch/deepcluster\" target=\"_blank\">https://github.com/facebookresearch/deepcluster</a></p>\n<p>Self-labelling via simultaneous clustering and representation learning<br>\n<a href=\"https://github.com/yukimasano/self-label\" target=\"_blank\">https://github.com/yukimasano/self-label</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976668,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "08/19/2020 02:56:16",
      "content": "<p>I believe that you can train KNeighbors with the kaggle gpu, which might be able to handle all of the data. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "976543": "I've essentially recreated the old Shazam algorithm that creates a \"constellation\" spectrogram, essentially a sparse matrix. \n\nI want to try fingerprinting all of the training data and then use a nearest neighbors approach for a baseline model. \n\nI'm concerned about how to go about loading so much data at once to fit sklearn's `KNeighborsClassifier`. Is this too ambitious? Sorry for the newb question. I'm new to handling such large datasets, as I'm recently coming from the beginner competitions like MNIST/Titanic/HomePrices. \n\nThanks!",
    "976668": "I believe that you can train KNeighbors with the kaggle gpu, which might be able to handle all of the data.",
    "976859": "I'm also a bit of a newbie - so disclaimer - the below may be wrong!\n\nGoing through this, I've tried to get features out of a semi-trained Resnet model (so just fed in the data to the model without really trying to make it accurate). So the data 'out' of this should have been somewhat representative of how 'similar' the calls were, some of their usual characteristics. I think it's usually termed embeddings, but I could be wrong.\n\nRunning TSNE on this was too tough - as each audio file has to have multiple entries for different time points - it wasn't really viable to fit everything into memory. Should probably have extracted less 'feature' data per row. Anyhow, I managed to get PCA to run (as at least you can fit on some of the data, then transform the rest). It did give some interesting starting points for further analysis, but it's worth knowing that although clips are labelled as 'bird X', in fact 'bird X' may have a whole variety of different calls, and there may be other birds singing as well. And I'm not even sure all the 'official' labels for every track are absolutely 100% correct - others who know more than me may have a view - it would probably not be surprising if there's a bit of a grey area after all this is very hard to identify correctly. But mostly I think the issue is other birds singing + variability of each bird type.\n\nJust mentioning this because maybe I went about it the wrong way but it seemed like a large and difficult task to extract some basic separating features across all clips. It did give me some useful info but was much messier than I'd anticipated.\n\nEdit - there is of course also just a lot of blank space, background noise in the audio files.",
    "977279": "Take a look at RAPIDS! Chris recently did some clustering on 2000+ dimensional embeddings of a few 10000's of images. It only took a few seconds :)",
    "977708": "Hey, I'll check it out, thanks! Someone else also recommended checking out Dask. Any thoughts on one vs the other (from the perspective of a relative beginner who joined the comp with less than a month left)?",
    "977717": "Dask is for distributed computing as far as I know while RAPIDS works entirely on GPU, so I think that's still orders of magnitude faster. Especially if you are using Kaggle Notebooks.",
    "977784": "Thank you, I appreciate the insight. Good luck to you!",
    "978096": "you may want to take a look at deep learning based clustering that could work quite well\n\nDeep Clustering for Unsupervised Learning of Visual Features\nhttps://github.com/facebookresearch/deepcluster\n\nSelf-labelling via simultaneous clustering and representation learning\nhttps://github.com/yukimasano/self-label",
    "992029": "And you can use RAPIDS with Dask!"
  },
  "source": "meta"
}