{
  "id": 176170,
  "title": "Clipping the training data",
  "url": "/competitions/birdsong-recognition/discussion/176170",
  "author_name": "",
  "post_date": "2020-08-20T18:40:27.269113700Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm looking for the best way to pre-process the training data. I've already resampled it all, but now I don't know what to do about the variable lengths.</p>\n<p>I see that the test data has to be labelled in 5-second increments and that it's a multi-label problem (I've never done that before). </p>\n<p>I also see that not every 5-second interval even has birdsong in it!</p>\n<p>I'm getting really confused as to how to proceed and would love some suggestions <strong>tailored to a relative ML beginner PLEASE</strong>. I've read a lot of other discussions, but they are way over my head. </p>",
  "messages": [
    {
      "id": "979294",
      "postDate": "08/20/2020 18:40:27",
      "content": "<p>I'm looking for the best way to pre-process the training data. I've already resampled it all, but now I don't know what to do about the variable lengths.</p>\n<p>I see that the test data has to be labelled in 5-second increments and that it's a multi-label problem (I've never done that before). </p>\n<p>I also see that not every 5-second interval even has birdsong in it!</p>\n<p>I'm getting really confused as to how to proceed and would love some suggestions <strong>tailored to a relative ML beginner PLEASE</strong>. I've read a lot of other discussions, but they are way over my head. </p>",
      "rawMarkdown": "I'm looking for the best way to pre-process the training data. I've already resampled it all, but now I don't know what to do about the variable lengths.\n\nI see that the test data has to be labelled in 5-second increments and that it's a multi-label problem (I've never done that before). \n\nI also see that not every 5-second interval even has birdsong in it!\n\nI'm getting really confused as to how to proceed and would love some suggestions **tailored to a relative ML beginner PLEASE**. I've read a lot of other discussions, but they are way over my head.",
      "votes": null
    },
    {
      "id": "979480",
      "postDate": "08/20/2020 21:58:42",
      "content": "<p>Hey, what I recommend is that you start by going through some of the public kernels \"Notebooks\" and seeing what they are doing. That is the best resource to understand how to begin coding for this competition.</p>",
      "rawMarkdown": "Hey, what I recommend is that you start by going through some of the public kernels \"Notebooks\" and seeing what they are doing. That is the best resource to understand how to begin coding for this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 979480,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "08/20/2020 21:58:42",
      "content": "<p>Hey, what I recommend is that you start by going through some of the public kernels \"Notebooks\" and seeing what they are doing. That is the best resource to understand how to begin coding for this competition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "979294": "I'm looking for the best way to pre-process the training data. I've already resampled it all, but now I don't know what to do about the variable lengths.\n\nI see that the test data has to be labelled in 5-second increments and that it's a multi-label problem (I've never done that before). \n\nI also see that not every 5-second interval even has birdsong in it!\n\nI'm getting really confused as to how to proceed and would love some suggestions **tailored to a relative ML beginner PLEASE**. I've read a lot of other discussions, but they are way over my head.",
    "979480": "Hey, what I recommend is that you start by going through some of the public kernels \"Notebooks\" and seeing what they are doing. That is the best resource to understand how to begin coding for this competition."
  },
  "source": "meta"
}