{
  "id": 220303,
  "title": "Template matching (couldn't get it to work)",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220303",
  "author_name": "",
  "post_date": "2021-02-18T00:00:43.966168300Z",
  "votes": 8,
  "comment_count": 3,
  "views": 0,
  "content": "<p>One of the things I keyed in on when I looked at this competition the first time was the puzzling way the data was labeled, with a separate tp and fp file. I read through the <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">paper</a> others had linked that was relevant to the way things had been labeled and noticed a couple key things. </p>\n<ul>\n<li>Data was initially labeled with template matching</li>\n<li>What was found via template matching was further reviewed by experts to put into tp and fp</li>\n</ul>\n<p>What I wanted to do with this info was instead of making the model detect based on the full audio or aggregated segments of the audio I would mimic the template matching procedure and have the model try to just copy the distinction made by the expert labelers.</p>\n<p>I was able to build a model that assumed I had the bounds of the time and freq that were output from the template matching procedure and yield validation of .985 lwlrap. This task is much easier, given two small spectrograms directly cropped on the point of interest which ones are the real species and which ones arent. Dont have to pray that your model learns from the relevant piece of a larger audio file and frequency ranges. So then the problem was figuring out how to derive the template matching procedure. If I could template match exactly like they template matched then a very high score was potentially achievable. </p>\n<p>The paper has a great amount of detail about the template matching procedure so I was able to even call the same functions that were used but was not quite able to match the process entirely. To explain the template matching procedure a bit, what they did is find exemplar audio segments that were great representations of the bird/frog noise based on signal to noise and expert choice. They drew a bounding box around the time and frequency. That is what is shown here:</p>\n<p><img src=\"https://ars.els-cdn.com/content/image/1-s2.0-S1574954120300637-gr2.jpg\" alt=\"\"> </p>\n<p>With these hand-selected examples of the various different species and the frequencies they occurred in they then did a 1D template matching procedure using <a href=\"https://scikit-image.org/docs/stable/auto_examples/features_detection/plot_template.html?highlight=match%20template\" target=\"_blank\">skimage match template</a> </p>\n<p>Confusing at first that it was only 1D template matching given that most of us were using 2D images in our spectrograms. What they did to simplify this to 1D is crop just to the frequency range of the original template. so if the frog was found in 4k-5k freq then they would crop 4k-5k of the spectrogram and then they only had to apply it across the time axis. This greatly simplifies the procedure of finding potential template matches because you don't have to worry as much about duplicates and small vertical shifts causing many false positives. Might have to apply non-max suppression or various other techniques to clean that up if not. </p>\n<p>So in the end there is a short maybe 30 x 500 pixel template that would be slide across a 30 x 10000 pixel long clip of audio and it would output a 1d signal showing how well the template matched across time with the clip of audio. This would output a plot like this: </p>\n<p><img src=\"https://i.imgur.com/CKDlRWg.png\" alt=\"\"></p>\n<p>With this signal they then used the <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks.html\" target=\"_blank\">find_peaks</a> function from scipy in order to find the probable candidates for the species given the correlation found from the template matching. This would look at the signal provided above and pick which time slices to actually propose as a likely match. A low threshold was used of .1 and matches within half the time span of the template could not be suggested. In theory, the output from this function, a list of time points, should yield all of the possible spots a species may have made noise. </p>\n<p>This was where I ran into issues and I'm not sure if others maybe used a technique like this and were able to implement it a bit better. No matter which was I cut it I always ended up with way too many false positives. In the paper they mention they found somewhere on the scale of 500k matches across many more audio segments than we had, but on our small set I was finding roughly that amount given the parameters they showed in the paper. </p>\n<p>One explanation for this was that I did not have the matching exemplar template that they located, I was just using the descendants from that initial clip, what was provided to us in the tp and fp files,  that only needed a match of .1 so my examples had potentially drifted quite a bit from the source golden sample. This is what some of them looked like for species 9:</p>\n<p><img src=\"https://i.imgur.com/oqtHLbe.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/tOLOdDK.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/tokhKTc.png\" alt=\"\"></p>\n<p>There is quite a bit of variation between these and lots of background noise it seems. If I tried to use each of these to find matches it would yield many many matches but they might not be very good and they would likely be redundant. Instead of this what I tried to do was take the average of all of these clips. My thought was that this would in some way recover a blurred version of the original template. The average would look something like this. The noise would be averaged away and hopefully, the signal would be left behind. </p>\n<p><img src=\"https://i.imgur.com/Uo502yJ.png\" alt=\"\"></p>\n<p>I tried this with just the tp, or the conjunction of the tp and fp, but regardless of which I used, I was seeing far too many false positives to really be useful. </p>\n<p>No satisfying conclusion to this just wanted to throw this out there to show what I attempted and discuss the technique I invested a bit of time into. </p>",
  "messages": [
    {
      "id": "1207588",
      "postDate": "02/18/2021 00:00:43",
      "content": "<p>One of the things I keyed in on when I looked at this competition the first time was the puzzling way the data was labeled, with a separate tp and fp file. I read through the <a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">paper</a> others had linked that was relevant to the way things had been labeled and noticed a couple key things. </p>\n<ul>\n<li>Data was initially labeled with template matching</li>\n<li>What was found via template matching was further reviewed by experts to put into tp and fp</li>\n</ul>\n<p>What I wanted to do with this info was instead of making the model detect based on the full audio or aggregated segments of the audio I would mimic the template matching procedure and have the model try to just copy the distinction made by the expert labelers.</p>\n<p>I was able to build a model that assumed I had the bounds of the time and freq that were output from the template matching procedure and yield validation of .985 lwlrap. This task is much easier, given two small spectrograms directly cropped on the point of interest which ones are the real species and which ones arent. Dont have to pray that your model learns from the relevant piece of a larger audio file and frequency ranges. So then the problem was figuring out how to derive the template matching procedure. If I could template match exactly like they template matched then a very high score was potentially achievable. </p>\n<p>The paper has a great amount of detail about the template matching procedure so I was able to even call the same functions that were used but was not quite able to match the process entirely. To explain the template matching procedure a bit, what they did is find exemplar audio segments that were great representations of the bird/frog noise based on signal to noise and expert choice. They drew a bounding box around the time and frequency. That is what is shown here:</p>\n<p><img src=\"https://ars.els-cdn.com/content/image/1-s2.0-S1574954120300637-gr2.jpg\" alt=\"\"> </p>\n<p>With these hand-selected examples of the various different species and the frequencies they occurred in they then did a 1D template matching procedure using <a href=\"https://scikit-image.org/docs/stable/auto_examples/features_detection/plot_template.html?highlight=match%20template\" target=\"_blank\">skimage match template</a> </p>\n<p>Confusing at first that it was only 1D template matching given that most of us were using 2D images in our spectrograms. What they did to simplify this to 1D is crop just to the frequency range of the original template. so if the frog was found in 4k-5k freq then they would crop 4k-5k of the spectrogram and then they only had to apply it across the time axis. This greatly simplifies the procedure of finding potential template matches because you don't have to worry as much about duplicates and small vertical shifts causing many false positives. Might have to apply non-max suppression or various other techniques to clean that up if not. </p>\n<p>So in the end there is a short maybe 30 x 500 pixel template that would be slide across a 30 x 10000 pixel long clip of audio and it would output a 1d signal showing how well the template matched across time with the clip of audio. This would output a plot like this: </p>\n<p><img src=\"https://i.imgur.com/CKDlRWg.png\" alt=\"\"></p>\n<p>With this signal they then used the <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks.html\" target=\"_blank\">find_peaks</a> function from scipy in order to find the probable candidates for the species given the correlation found from the template matching. This would look at the signal provided above and pick which time slices to actually propose as a likely match. A low threshold was used of .1 and matches within half the time span of the template could not be suggested. In theory, the output from this function, a list of time points, should yield all of the possible spots a species may have made noise. </p>\n<p>This was where I ran into issues and I'm not sure if others maybe used a technique like this and were able to implement it a bit better. No matter which was I cut it I always ended up with way too many false positives. In the paper they mention they found somewhere on the scale of 500k matches across many more audio segments than we had, but on our small set I was finding roughly that amount given the parameters they showed in the paper. </p>\n<p>One explanation for this was that I did not have the matching exemplar template that they located, I was just using the descendants from that initial clip, what was provided to us in the tp and fp files,  that only needed a match of .1 so my examples had potentially drifted quite a bit from the source golden sample. This is what some of them looked like for species 9:</p>\n<p><img src=\"https://i.imgur.com/oqtHLbe.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/tOLOdDK.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/tokhKTc.png\" alt=\"\"></p>\n<p>There is quite a bit of variation between these and lots of background noise it seems. If I tried to use each of these to find matches it would yield many many matches but they might not be very good and they would likely be redundant. Instead of this what I tried to do was take the average of all of these clips. My thought was that this would in some way recover a blurred version of the original template. The average would look something like this. The noise would be averaged away and hopefully, the signal would be left behind. </p>\n<p><img src=\"https://i.imgur.com/Uo502yJ.png\" alt=\"\"></p>\n<p>I tried this with just the tp, or the conjunction of the tp and fp, but regardless of which I used, I was seeing far too many false positives to really be useful. </p>\n<p>No satisfying conclusion to this just wanted to throw this out there to show what I attempted and discuss the technique I invested a bit of time into. </p>",
      "rawMarkdown": "One of the things I keyed in on when I looked at this competition the first time was the puzzling way the data was labeled, with a separate tp and fp file. I read through the [paper](https://www.sciencedirect.com/science/article/pii/S1574954120300637) others had linked that was relevant to the way things had been labeled and noticed a couple key things. \n\n- Data was initially labeled with template matching\n- What was found via template matching was further reviewed by experts to put into tp and fp\n\nWhat I wanted to do with this info was instead of making the model detect based on the full audio or aggregated segments of the audio I would mimic the template matching procedure and have the model try to just copy the distinction made by the expert labelers.\n\nI was able to build a model that assumed I had the bounds of the time and freq that were output from the template matching procedure and yield validation of .985 lwlrap. This task is much easier, given two small spectrograms directly cropped on the point of interest which ones are the real species and which ones arent. Dont have to pray that your model learns from the relevant piece of a larger audio file and frequency ranges. So then the problem was figuring out how to derive the template matching procedure. If I could template match exactly like they template matched then a very high score was potentially achievable. \n\nThe paper has a great amount of detail about the template matching procedure so I was able to even call the same functions that were used but was not quite able to match the process entirely. To explain the template matching procedure a bit, what they did is find exemplar audio segments that were great representations of the bird/frog noise based on signal to noise and expert choice. They drew a bounding box around the time and frequency. That is what is shown here:\n\n![](https://ars.els-cdn.com/content/image/1-s2.0-S1574954120300637-gr2.jpg) \n\nWith these hand-selected examples of the various different species and the frequencies they occurred in they then did a 1D template matching procedure using [skimage match template](https://scikit-image.org/docs/stable/auto_examples/features_detection/plot_template.html?highlight=match%20template) \n\nConfusing at first that it was only 1D template matching given that most of us were using 2D images in our spectrograms. What they did to simplify this to 1D is crop just to the frequency range of the original template. so if the frog was found in 4k-5k freq then they would crop 4k-5k of the spectrogram and then they only had to apply it across the time axis. This greatly simplifies the procedure of finding potential template matches because you don't have to worry as much about duplicates and small vertical shifts causing many false positives. Might have to apply non-max suppression or various other techniques to clean that up if not. \n\nSo in the end there is a short maybe 30 x 500 pixel template that would be slide across a 30 x 10000 pixel long clip of audio and it would output a 1d signal showing how well the template matched across time with the clip of audio. This would output a plot like this: \n\n![](https://i.imgur.com/CKDlRWg.png)\n\nWith this signal they then used the [find_peaks](https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks.html) function from scipy in order to find the probable candidates for the species given the correlation found from the template matching. This would look at the signal provided above and pick which time slices to actually propose as a likely match. A low threshold was used of .1 and matches within half the time span of the template could not be suggested. In theory, the output from this function, a list of time points, should yield all of the possible spots a species may have made noise. \n\nThis was where I ran into issues and I'm not sure if others maybe used a technique like this and were able to implement it a bit better. No matter which was I cut it I always ended up with way too many false positives. In the paper they mention they found somewhere on the scale of 500k matches across many more audio segments than we had, but on our small set I was finding roughly that amount given the parameters they showed in the paper. \n\nOne explanation for this was that I did not have the matching exemplar template that they located, I was just using the descendants from that initial clip, what was provided to us in the tp and fp files,  that only needed a match of .1 so my examples had potentially drifted quite a bit from the source golden sample. This is what some of them looked like for species 9:\n\n \n![](https://i.imgur.com/oqtHLbe.png)\n\n![](https://i.imgur.com/tOLOdDK.png)\n\n![](https://i.imgur.com/tokhKTc.png)\n\nThere is quite a bit of variation between these and lots of background noise it seems. If I tried to use each of these to find matches it would yield many many matches but they might not be very good and they would likely be redundant. Instead of this what I tried to do was take the average of all of these clips. My thought was that this would in some way recover a blurred version of the original template. The average would look something like this. The noise would be averaged away and hopefully, the signal would be left behind. \n\n![](https://i.imgur.com/Uo502yJ.png)\n\nI tried this with just the tp, or the conjunction of the tp and fp, but regardless of which I used, I was seeing far too many false positives to really be useful. \n\nNo satisfying conclusion to this just wanted to throw this out there to show what I attempted and discuss the technique I invested a bit of time into.",
      "votes": null
    },
    {
      "id": "1207636",
      "postDate": "02/18/2021 00:30:46",
      "content": "<p>the template is a 2d template in the spectrogram.<br>\nmake a mean image of the annotation, you will see this:<br>\n <img src=\"https://i.ibb.co/hMFK1zm/11-1-box.png\" alt=\"\"></p>\n<p>you can then use the mean as a template to find rois that is similar to the mean.</p>\n<p>but template matching is essentially a conv, so you might as well train a  conv classifier</p>",
      "rawMarkdown": "the template is a 2d template in the spectrogram.\nmake a mean image of the annotation, you will see this:\n ![](https://i.ibb.co/hMFK1zm/11-1-box.png)\n\n\nyou can then use the mean as a template to find rois that is similar to the mean.\n\nbut template matching is essentially a conv, so you might as well train a  conv classifier",
      "votes": null
    },
    {
      "id": "1207693",
      "postDate": "02/18/2021 01:18:49",
      "content": "<p>Yes, I did take the mean of the labeled regions to find the template, but yes you might as well train a cnn to find the same thing</p>",
      "rawMarkdown": "Yes, I did take the mean of the labeled regions to find the template, but yes you might as well train a cnn to find the same thing",
      "votes": null
    },
    {
      "id": "2327938",
      "postDate": "07/03/2023 08:55:39",
      "content": "<p>Realtime template matching by SAM model: template matching using SAM model : <a href=\"https://youtu.be/ufoithWSv4U\" target=\"_blank\">https://youtu.be/ufoithWSv4U</a></p>",
      "rawMarkdown": "Realtime template matching by SAM model: template matching using SAM model : https://youtu.be/ufoithWSv4U",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207636,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/18/2021 00:30:46",
      "content": "<p>the template is a 2d template in the spectrogram.<br>\nmake a mean image of the annotation, you will see this:<br>\n <img src=\"https://i.ibb.co/hMFK1zm/11-1-box.png\" alt=\"\"></p>\n<p>you can then use the mean as a template to find rois that is similar to the mean.</p>\n<p>but template matching is essentially a conv, so you might as well train a  conv classifier</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207693,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/18/2021 01:18:49",
          "content": "<p>Yes, I did take the mean of the labeled regions to find the template, but yes you might as well train a cnn to find the same thing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2327938,
      "author_name": "bemorekgg",
      "author_url": "",
      "post_date": "07/03/2023 08:55:39",
      "content": "<p>Realtime template matching by SAM model: template matching using SAM model : <a href=\"https://youtu.be/ufoithWSv4U\" target=\"_blank\">https://youtu.be/ufoithWSv4U</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207588": "One of the things I keyed in on when I looked at this competition the first time was the puzzling way the data was labeled, with a separate tp and fp file. I read through the [paper](https://www.sciencedirect.com/science/article/pii/S1574954120300637) others had linked that was relevant to the way things had been labeled and noticed a couple key things. \n\n- Data was initially labeled with template matching\n- What was found via template matching was further reviewed by experts to put into tp and fp\n\nWhat I wanted to do with this info was instead of making the model detect based on the full audio or aggregated segments of the audio I would mimic the template matching procedure and have the model try to just copy the distinction made by the expert labelers.\n\nI was able to build a model that assumed I had the bounds of the time and freq that were output from the template matching procedure and yield validation of .985 lwlrap. This task is much easier, given two small spectrograms directly cropped on the point of interest which ones are the real species and which ones arent. Dont have to pray that your model learns from the relevant piece of a larger audio file and frequency ranges. So then the problem was figuring out how to derive the template matching procedure. If I could template match exactly like they template matched then a very high score was potentially achievable. \n\nThe paper has a great amount of detail about the template matching procedure so I was able to even call the same functions that were used but was not quite able to match the process entirely. To explain the template matching procedure a bit, what they did is find exemplar audio segments that were great representations of the bird/frog noise based on signal to noise and expert choice. They drew a bounding box around the time and frequency. That is what is shown here:\n\n![](https://ars.els-cdn.com/content/image/1-s2.0-S1574954120300637-gr2.jpg) \n\nWith these hand-selected examples of the various different species and the frequencies they occurred in they then did a 1D template matching procedure using [skimage match template](https://scikit-image.org/docs/stable/auto_examples/features_detection/plot_template.html?highlight=match%20template) \n\nConfusing at first that it was only 1D template matching given that most of us were using 2D images in our spectrograms. What they did to simplify this to 1D is crop just to the frequency range of the original template. so if the frog was found in 4k-5k freq then they would crop 4k-5k of the spectrogram and then they only had to apply it across the time axis. This greatly simplifies the procedure of finding potential template matches because you don't have to worry as much about duplicates and small vertical shifts causing many false positives. Might have to apply non-max suppression or various other techniques to clean that up if not. \n\nSo in the end there is a short maybe 30 x 500 pixel template that would be slide across a 30 x 10000 pixel long clip of audio and it would output a 1d signal showing how well the template matched across time with the clip of audio. This would output a plot like this: \n\n![](https://i.imgur.com/CKDlRWg.png)\n\nWith this signal they then used the [find_peaks](https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.find_peaks.html) function from scipy in order to find the probable candidates for the species given the correlation found from the template matching. This would look at the signal provided above and pick which time slices to actually propose as a likely match. A low threshold was used of .1 and matches within half the time span of the template could not be suggested. In theory, the output from this function, a list of time points, should yield all of the possible spots a species may have made noise. \n\nThis was where I ran into issues and I'm not sure if others maybe used a technique like this and were able to implement it a bit better. No matter which was I cut it I always ended up with way too many false positives. In the paper they mention they found somewhere on the scale of 500k matches across many more audio segments than we had, but on our small set I was finding roughly that amount given the parameters they showed in the paper. \n\nOne explanation for this was that I did not have the matching exemplar template that they located, I was just using the descendants from that initial clip, what was provided to us in the tp and fp files,  that only needed a match of .1 so my examples had potentially drifted quite a bit from the source golden sample. This is what some of them looked like for species 9:\n\n \n![](https://i.imgur.com/oqtHLbe.png)\n\n![](https://i.imgur.com/tOLOdDK.png)\n\n![](https://i.imgur.com/tokhKTc.png)\n\nThere is quite a bit of variation between these and lots of background noise it seems. If I tried to use each of these to find matches it would yield many many matches but they might not be very good and they would likely be redundant. Instead of this what I tried to do was take the average of all of these clips. My thought was that this would in some way recover a blurred version of the original template. The average would look something like this. The noise would be averaged away and hopefully, the signal would be left behind. \n\n![](https://i.imgur.com/Uo502yJ.png)\n\nI tried this with just the tp, or the conjunction of the tp and fp, but regardless of which I used, I was seeing far too many false positives to really be useful. \n\nNo satisfying conclusion to this just wanted to throw this out there to show what I attempted and discuss the technique I invested a bit of time into.",
    "1207636": "the template is a 2d template in the spectrogram.\nmake a mean image of the annotation, you will see this:\n ![](https://i.ibb.co/hMFK1zm/11-1-box.png)\n\n\nyou can then use the mean as a template to find rois that is similar to the mean.\n\nbut template matching is essentially a conv, so you might as well train a  conv classifier",
    "1207693": "Yes, I did take the mean of the labeled regions to find the template, but yes you might as well train a cnn to find the same thing",
    "2327938": "Realtime template matching by SAM model: template matching using SAM model : https://youtu.be/ufoithWSv4U"
  },
  "source": "meta"
}