{
  "id": 217166,
  "title": "Beginner Question: How to Apply my Model to the Test Data",
  "url": "/competitions/rfcx-species-audio-detection/discussion/217166",
  "author_name": "",
  "post_date": "2021-02-05T15:30:31.076937Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This is the first real attempt at Kaggle and I'm stuck on a very basic, high level issue with this competition. This is: how to apply what my model learns to this particular test set, and how to get the correct type of submission results.</p>\n<p>Possibly I'm thinking about this the wrong way, any help would be <em>greatly appreciated</em>. I'm finding so much documentation about how audio is processed and how CNNs work, etc. But I am having trouble finding advice on this very basic question.</p>\n<p>Let's say this is my approach: </p>\n<p><strong>Processing my Training Data:</strong></p>\n<p>Using the start and end times in the train_tp file, I create a bunch of six second audio clips. I do this by finding the middle point of each t_min and t_max, and centering a six-second clip on that. I then create a mel spectrogram from each six-second audio clip, and save it with the correct label.</p>\n<p><strong>Creating a Model:</strong></p>\n<p>Using the data above, I create a model that can predict a label based on the mel spectrogram for a six-second audio clip.</p>\n<p><strong>Applying my Model to Test Data:</strong></p>\n<p>Here is where my (2) questions lie:</p>\n<p>(1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?</p>\n<p>(2) My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?</p>\n<p>Thanks in advance for any advice.</p>",
  "messages": [
    {
      "id": "1187625",
      "postDate": "02/05/2021 15:30:31",
      "content": "<p>This is the first real attempt at Kaggle and I'm stuck on a very basic, high level issue with this competition. This is: how to apply what my model learns to this particular test set, and how to get the correct type of submission results.</p>\n<p>Possibly I'm thinking about this the wrong way, any help would be <em>greatly appreciated</em>. I'm finding so much documentation about how audio is processed and how CNNs work, etc. But I am having trouble finding advice on this very basic question.</p>\n<p>Let's say this is my approach: </p>\n<p><strong>Processing my Training Data:</strong></p>\n<p>Using the start and end times in the train_tp file, I create a bunch of six second audio clips. I do this by finding the middle point of each t_min and t_max, and centering a six-second clip on that. I then create a mel spectrogram from each six-second audio clip, and save it with the correct label.</p>\n<p><strong>Creating a Model:</strong></p>\n<p>Using the data above, I create a model that can predict a label based on the mel spectrogram for a six-second audio clip.</p>\n<p><strong>Applying my Model to Test Data:</strong></p>\n<p>Here is where my (2) questions lie:</p>\n<p>(1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?</p>\n<p>(2) My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?</p>\n<p>Thanks in advance for any advice.</p>",
      "rawMarkdown": "This is the first real attempt at Kaggle and I'm stuck on a very basic, high level issue with this competition. This is: how to apply what my model learns to this particular test set, and how to get the correct type of submission results.\n\nPossibly I'm thinking about this the wrong way, any help would be *greatly appreciated*. I'm finding so much documentation about how audio is processed and how CNNs work, etc. But I am having trouble finding advice on this very basic question.\n\nLet's say this is my approach: \n\n**Processing my Training Data:**\n\nUsing the start and end times in the train_tp file, I create a bunch of six second audio clips. I do this by finding the middle point of each t_min and t_max, and centering a six-second clip on that. I then create a mel spectrogram from each six-second audio clip, and save it with the correct label.\n\n\n**Creating a Model:**\n\nUsing the data above, I create a model that can predict a label based on the mel spectrogram for a six-second audio clip.\n\n**Applying my Model to Test Data:**\n\nHere is where my (2) questions lie:\n\n(1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?\n\n(2) My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?\n\nThanks in advance for any advice.",
      "votes": null
    },
    {
      "id": "1187689",
      "postDate": "02/05/2021 16:26:52",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/benbray1\" target=\"_blank\">@benbray1</a> , welcome and hope you enjoy 🤓 </p>\n<p>There was a similar topic <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212863\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?</p>\n</blockquote>\n<p>I think the simplest logic is that, if a bird exists in any of your 6 second clips, then, the bird exists in this 60s audio, right? So, maybe you can use the max of your predictions of 6 second clips as your prediction of this 60s audio.</p>",
      "rawMarkdown": "Hi, @benbray1 , welcome and hope you enjoy 🤓 \n\nThere was a similar topic [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212863).\n\n> My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?\n\nI think the simplest logic is that, if a bird exists in any of your 6 second clips, then, the bird exists in this 60s audio, right? So, maybe you can use the max of your predictions of 6 second clips as your prediction of this 60s audio.",
      "votes": null
    },
    {
      "id": "1187723",
      "postDate": "02/05/2021 16:56:47",
      "content": "<p>Okay great, that's enough to get me started. I was wondering if I was supposed to add them up and normalize them, or what. But using max makes more sense. I will try that. Thank you <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> ! </p>",
      "rawMarkdown": "Okay great, that's enough to get me started. I was wondering if I was supposed to add them up and normalize them, or what. But using max makes more sense. I will try that. Thank you @barnwellguy !",
      "votes": null
    },
    {
      "id": "1188004",
      "postDate": "02/05/2021 21:35:45",
      "content": "<p>Entering this as your first Kaggle competition is courageous.  It is a hard one, because train is only partly labelled and test is quite different.</p>\n<blockquote>\n  <p>(1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?</p>\n</blockquote>\n<p>I always give the same answer to this kind of question: try both and see what works best.</p>\n<p><a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> answered the second question as I would have ;)</p>",
      "rawMarkdown": "Entering this as your first Kaggle competition is courageous.  It is a hard one, because train is only partly labelled and test is quite different.\n\n> (1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?\n\nI always give the same answer to this kind of question: try both and see what works best.\n\n@barnwellguy answered the second question as I would have ;)",
      "votes": null
    },
    {
      "id": "1188995",
      "postDate": "02/06/2021 16:28:57",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . That context is encouraging. It definitely felt like there were a few curveballs in this problem , but I wasn't sure :)</p>",
      "rawMarkdown": "Thank you @cpmpml . That context is encouraging. It definitely felt like there were a few curveballs in this problem , but I wasn't sure :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1187689,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "02/05/2021 16:26:52",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/benbray1\" target=\"_blank\">@benbray1</a> , welcome and hope you enjoy 🤓 </p>\n<p>There was a similar topic <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212863\" target=\"_blank\">here</a>.</p>\n<blockquote>\n  <p>My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?</p>\n</blockquote>\n<p>I think the simplest logic is that, if a bird exists in any of your 6 second clips, then, the bird exists in this 60s audio, right? So, maybe you can use the max of your predictions of 6 second clips as your prediction of this 60s audio.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1187723,
      "author_name": "benbray1",
      "author_url": "",
      "post_date": "02/05/2021 16:56:47",
      "content": "<p>Okay great, that's enough to get me started. I was wondering if I was supposed to add them up and normalize them, or what. But using max makes more sense. I will try that. Thank you <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1188004,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/05/2021 21:35:45",
      "content": "<p>Entering this as your first Kaggle competition is courageous.  It is a hard one, because train is only partly labelled and test is quite different.</p>\n<blockquote>\n  <p>(1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?</p>\n</blockquote>\n<p>I always give the same answer to this kind of question: try both and see what works best.</p>\n<p><a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> answered the second question as I would have ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1188995,
      "author_name": "benbray1",
      "author_url": "",
      "post_date": "02/06/2021 16:28:57",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> . That context is encouraging. It definitely felt like there were a few curveballs in this problem , but I wasn't sure :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1187625": "This is the first real attempt at Kaggle and I'm stuck on a very basic, high level issue with this competition. This is: how to apply what my model learns to this particular test set, and how to get the correct type of submission results.\n\nPossibly I'm thinking about this the wrong way, any help would be *greatly appreciated*. I'm finding so much documentation about how audio is processed and how CNNs work, etc. But I am having trouble finding advice on this very basic question.\n\nLet's say this is my approach: \n\n**Processing my Training Data:**\n\nUsing the start and end times in the train_tp file, I create a bunch of six second audio clips. I do this by finding the middle point of each t_min and t_max, and centering a six-second clip on that. I then create a mel spectrogram from each six-second audio clip, and save it with the correct label.\n\n\n**Creating a Model:**\n\nUsing the data above, I create a model that can predict a label based on the mel spectrogram for a six-second audio clip.\n\n**Applying my Model to Test Data:**\n\nHere is where my (2) questions lie:\n\n(1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?\n\n(2) My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?\n\nThanks in advance for any advice.",
    "1187689": "Hi, @benbray1 , welcome and hope you enjoy 🤓 \n\nThere was a similar topic [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212863).\n\n> My model predicts a label for a six second clip, but we need to submit answers for each 60 second file. How do I combine my predictions from all my shorter clips into an answer for the whole file?\n\nI think the simplest logic is that, if a bird exists in any of your 6 second clips, then, the bird exists in this 60s audio, right? So, maybe you can use the max of your predictions of 6 second clips as your prediction of this 60s audio.",
    "1187723": "Okay great, that's enough to get me started. I was wondering if I was supposed to add them up and normalize them, or what. But using max makes more sense. I will try that. Thank you @barnwellguy !",
    "1188004": "Entering this as your first Kaggle competition is courageous.  It is a hard one, because train is only partly labelled and test is quite different.\n\n> (1) My model can deal with six second clips. So I assume I first have to split up the test audio. Should I split each file into 10 six-second clips? Or should there be more clips that overlap?\n\nI always give the same answer to this kind of question: try both and see what works best.\n\n@barnwellguy answered the second question as I would have ;)",
    "1188995": "Thank you @cpmpml . That context is encouraging. It definitely felt like there were a few curveballs in this problem , but I wasn't sure :)"
  },
  "source": "meta"
}