{
  "id": 220302,
  "title": "Visualizing predictions over time",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220302",
  "author_name": "",
  "post_date": "2021-02-17T23:57:44.439785600Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>A small diagnosing step I started doing for my models was looking at predictions over time like from this <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">post</a>. </p>\n<p>Most people are doing some sort of sliding window prediction, like predicting on 10s of audio then moving the period forward by 5s and predicting again. Because of this we have a time-series plot was can make of our predictions through the successive frames. Looking at this was pretty interesting to me because I would think that we would have relatively low output and then spikes to represent when the model thought it heard a given species. I thought that some of my error may have been coming from predictions away from the audio of interest, a false positive occurring outside of the region where the bird/frog actually made a sound, but what I found was significantly different than that. </p>\n<p>This kind of points to the model picking up on background audio or some other signal in the image rather than the bird/frog audio itself. This is the output from a model in which I made predictions for the different frequency bands for the different species and then had it predict each individually. Much less clear than I would have expected. Does not make sense that the predictions are so constantly high for many species. </p>\n<p>Like others have pointed out it seemed like s3 was consistently problematic but even other species did not behave how I expected. Curious if others found similar patterns in their model's output or if that was some failure of my own. </p>\n<p><img src=\"https://i.imgur.com/VnYbPLu.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1207586",
      "postDate": "02/17/2021 23:57:44",
      "content": "<p>A small diagnosing step I started doing for my models was looking at predictions over time like from this <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">post</a>. </p>\n<p>Most people are doing some sort of sliding window prediction, like predicting on 10s of audio then moving the period forward by 5s and predicting again. Because of this we have a time-series plot was can make of our predictions through the successive frames. Looking at this was pretty interesting to me because I would think that we would have relatively low output and then spikes to represent when the model thought it heard a given species. I thought that some of my error may have been coming from predictions away from the audio of interest, a false positive occurring outside of the region where the bird/frog actually made a sound, but what I found was significantly different than that. </p>\n<p>This kind of points to the model picking up on background audio or some other signal in the image rather than the bird/frog audio itself. This is the output from a model in which I made predictions for the different frequency bands for the different species and then had it predict each individually. Much less clear than I would have expected. Does not make sense that the predictions are so constantly high for many species. </p>\n<p>Like others have pointed out it seemed like s3 was consistently problematic but even other species did not behave how I expected. Curious if others found similar patterns in their model's output or if that was some failure of my own. </p>\n<p><img src=\"https://i.imgur.com/VnYbPLu.png\" alt=\"\"></p>",
      "rawMarkdown": "A small diagnosing step I started doing for my models was looking at predictions over time like from this [post](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040). \n\nMost people are doing some sort of sliding window prediction, like predicting on 10s of audio then moving the period forward by 5s and predicting again. Because of this we have a time-series plot was can make of our predictions through the successive frames. Looking at this was pretty interesting to me because I would think that we would have relatively low output and then spikes to represent when the model thought it heard a given species. I thought that some of my error may have been coming from predictions away from the audio of interest, a false positive occurring outside of the region where the bird/frog actually made a sound, but what I found was significantly different than that. \n\nThis kind of points to the model picking up on background audio or some other signal in the image rather than the bird/frog audio itself. This is the output from a model in which I made predictions for the different frequency bands for the different species and then had it predict each individually. Much less clear than I would have expected. Does not make sense that the predictions are so constantly high for many species. \n\nLike others have pointed out it seemed like s3 was consistently problematic but even other species did not behave how I expected. Curious if others found similar patterns in their model's output or if that was some failure of my own. \n\n![](https://i.imgur.com/VnYbPLu.png)",
      "votes": null
    },
    {
      "id": "1207625",
      "postDate": "02/18/2021 00:24:15",
      "content": "<p>I was curious on same thing when trying to evaluate different inference schemes.. Could you please confirm what are your axis x,y in the figure above so I can reproduce smth similar to compare ? x-axis are the splited audio clips &amp; y-axis probs (after sigmoid) maybe ? </p>",
      "rawMarkdown": "I was curious on same thing when trying to evaluate different inference schemes.. Could you please confirm what are your axis x,y in the figure above so I can reproduce smth similar to compare ? x-axis are the splited audio clips & y-axis probs (after sigmoid) maybe ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207625,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "02/18/2021 00:24:15",
      "content": "<p>I was curious on same thing when trying to evaluate different inference schemes.. Could you please confirm what are your axis x,y in the figure above so I can reproduce smth similar to compare ? x-axis are the splited audio clips &amp; y-axis probs (after sigmoid) maybe ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207586": "A small diagnosing step I started doing for my models was looking at predictions over time like from this [post](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040). \n\nMost people are doing some sort of sliding window prediction, like predicting on 10s of audio then moving the period forward by 5s and predicting again. Because of this we have a time-series plot was can make of our predictions through the successive frames. Looking at this was pretty interesting to me because I would think that we would have relatively low output and then spikes to represent when the model thought it heard a given species. I thought that some of my error may have been coming from predictions away from the audio of interest, a false positive occurring outside of the region where the bird/frog actually made a sound, but what I found was significantly different than that. \n\nThis kind of points to the model picking up on background audio or some other signal in the image rather than the bird/frog audio itself. This is the output from a model in which I made predictions for the different frequency bands for the different species and then had it predict each individually. Much less clear than I would have expected. Does not make sense that the predictions are so constantly high for many species. \n\nLike others have pointed out it seemed like s3 was consistently problematic but even other species did not behave how I expected. Curious if others found similar patterns in their model's output or if that was some failure of my own. \n\n![](https://i.imgur.com/VnYbPLu.png)",
    "1207625": "I was curious on same thing when trying to evaluate different inference schemes.. Could you please confirm what are your axis x,y in the figure above so I can reproduce smth similar to compare ? x-axis are the splited audio clips & y-axis probs (after sigmoid) maybe ?"
  },
  "source": "meta"
}