{
  "id": 201000,
  "title": "Data Transformations",
  "url": "/competitions/rfcx-species-audio-detection/discussion/201000",
  "author_name": "",
  "post_date": "2020-12-02T17:49:53.504346500Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>What kind of data transformations are you using? I'm looking at the following transformations:</p>\n<ul>\n<li>Frequency masking</li>\n<li>Time masking</li>\n<li>White, pink noise injection</li>\n<li>Pitch shifting<ul>\n<li>Only between Tmin and Tmax</li>\n<li>Only shifting frequencies between Fmin and Fmax, and shifting them within a 95% confidence interval of all the Fmins and Fmaxes for that specific species_id</li></ul></li>\n<li>Selecting random N second crops of the audio, where N is determined species by species<ul>\n<li>Using a 10 second crop may be too large for songs that only last 1 second on average, for example.</li></ul></li>\n</ul>\n<p>Feel free to add to the list!</p>",
  "messages": [
    {
      "id": "1099939",
      "postDate": "12/02/2020 17:49:53",
      "content": "<p>What kind of data transformations are you using? I'm looking at the following transformations:</p>\n<ul>\n<li>Frequency masking</li>\n<li>Time masking</li>\n<li>White, pink noise injection</li>\n<li>Pitch shifting<ul>\n<li>Only between Tmin and Tmax</li>\n<li>Only shifting frequencies between Fmin and Fmax, and shifting them within a 95% confidence interval of all the Fmins and Fmaxes for that specific species_id</li></ul></li>\n<li>Selecting random N second crops of the audio, where N is determined species by species<ul>\n<li>Using a 10 second crop may be too large for songs that only last 1 second on average, for example.</li></ul></li>\n</ul>\n<p>Feel free to add to the list!</p>",
      "rawMarkdown": "What kind of data transformations are you using? I'm looking at the following transformations:\n- Frequency masking\n- Time masking\n- White, pink noise injection\n- Pitch shifting\n     - Only between Tmin and Tmax\n     - Only shifting frequencies between Fmin and Fmax, and shifting them within a 95% confidence interval of all the Fmins and Fmaxes for that specific species_id\n- Selecting random N second crops of the audio, where N is determined species by species\n     - Using a 10 second crop may be too large for songs that only last 1 second on average, for example.\n\nFeel free to add to the list!",
      "votes": null
    },
    {
      "id": "1104250",
      "postDate": "12/06/2020 18:46:49",
      "content": "<p><a href=\"https://www.kaggle.com/maltonji\" target=\"_blank\">@maltonji</a> Some good ideas. But when I checked the CSV and tried to find sequences, I have noticed that t_min and t_max is not accurate so your last idea might not work as good as intended.</p>",
      "rawMarkdown": "maltonji Some good ideas. But when I checked the CSV and tried to find sequences, I have noticed that t_min and t_max is not accurate so your last idea might not work as good as intended.",
      "votes": null
    },
    {
      "id": "1104353",
      "postDate": "12/06/2020 20:56:22",
      "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> From looking at the csv, it looks like a lot of the data has the exact same time duration. Is that what you mean? </p>\n<p>Because there are 2 files, one for True Positives and one for False Positives, it's implied that they've been reviewed for accuracy. Therefore, I'd trust that t_min and t_max are at least **close **to where the song is identified. Still, I'd agree the duration could be off to some degree.</p>",
      "rawMarkdown": "aliabdin1 From looking at the csv, it looks like a lot of the data has the exact same time duration. Is that what you mean? \n\nBecause there are 2 files, one for True Positives and one for False Positives, it's implied that they've been reviewed for accuracy. Therefore, I'd trust that t_min and t_max are at least **close **to where the song is identified. Still, I'd agree the duration could be off to some degree.",
      "votes": null
    },
    {
      "id": "1104373",
      "postDate": "12/06/2020 21:34:24",
      "content": "<p>I also mean the time t_min and t_max, I took random audio_files and checked them if I could hear anything during the time and sometimes there was no noise between the interval [t_min, t_max]</p>",
      "rawMarkdown": "I also mean the time t_min and t_max, I took random audio_files and checked them if I could hear anything during the time and sometimes there was no noise between the interval [t_min, t_max]",
      "votes": null
    },
    {
      "id": "1104378",
      "postDate": "12/06/2020 21:44:13",
      "content": "<p>Unless t_min == t_max, you should hear audio. Since it comes from a continuous 60-second audio clip, there will always be sound (even if it's just background noise). Just a thought!</p>",
      "rawMarkdown": "Unless t_min == t_max, you should hear audio. Since it comes from a continuous 60-second audio clip, there will always be sound (even if it's just background noise). Just a thought!",
      "votes": null
    },
    {
      "id": "1104382",
      "postDate": "12/06/2020 21:55:00",
      "content": "<p>I do hear audio, but I dont hear the species which should be making noise during this interval.</p>",
      "rawMarkdown": "I do hear audio, but I dont hear the species which should be making noise during this interval.",
      "votes": null
    },
    {
      "id": "1109674",
      "postDate": "12/12/2020 00:15:48",
      "content": "<p>Another augmentation method for audio is reverberation</p>",
      "rawMarkdown": "Another augmentation method for audio is reverberation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1104250,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "12/06/2020 18:46:49",
      "content": "<p><a href=\"https://www.kaggle.com/maltonji\" target=\"_blank\">@maltonji</a> Some good ideas. But when I checked the CSV and tried to find sequences, I have noticed that t_min and t_max is not accurate so your last idea might not work as good as intended.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1104353,
          "author_name": "maltonji",
          "author_url": "",
          "post_date": "12/06/2020 20:56:22",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> From looking at the csv, it looks like a lot of the data has the exact same time duration. Is that what you mean? </p>\n<p>Because there are 2 files, one for True Positives and one for False Positives, it's implied that they've been reviewed for accuracy. Therefore, I'd trust that t_min and t_max are at least **close **to where the song is identified. Still, I'd agree the duration could be off to some degree.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104373,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "12/06/2020 21:34:24",
          "content": "<p>I also mean the time t_min and t_max, I took random audio_files and checked them if I could hear anything during the time and sometimes there was no noise between the interval [t_min, t_max]</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104378,
          "author_name": "maltonji",
          "author_url": "",
          "post_date": "12/06/2020 21:44:13",
          "content": "<p>Unless t_min == t_max, you should hear audio. Since it comes from a continuous 60-second audio clip, there will always be sound (even if it's just background noise). Just a thought!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104382,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "12/06/2020 21:55:00",
          "content": "<p>I do hear audio, but I dont hear the species which should be making noise during this interval.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109674,
      "author_name": "olmatz",
      "author_url": "",
      "post_date": "12/12/2020 00:15:48",
      "content": "<p>Another augmentation method for audio is reverberation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1099939": "What kind of data transformations are you using? I'm looking at the following transformations:\n- Frequency masking\n- Time masking\n- White, pink noise injection\n- Pitch shifting\n     - Only between Tmin and Tmax\n     - Only shifting frequencies between Fmin and Fmax, and shifting them within a 95% confidence interval of all the Fmins and Fmaxes for that specific species_id\n- Selecting random N second crops of the audio, where N is determined species by species\n     - Using a 10 second crop may be too large for songs that only last 1 second on average, for example.\n\nFeel free to add to the list!",
    "1104250": "maltonji Some good ideas. But when I checked the CSV and tried to find sequences, I have noticed that t_min and t_max is not accurate so your last idea might not work as good as intended.",
    "1104353": "aliabdin1 From looking at the csv, it looks like a lot of the data has the exact same time duration. Is that what you mean? \n\nBecause there are 2 files, one for True Positives and one for False Positives, it's implied that they've been reviewed for accuracy. Therefore, I'd trust that t_min and t_max are at least **close **to where the song is identified. Still, I'd agree the duration could be off to some degree.",
    "1104373": "I also mean the time t_min and t_max, I took random audio_files and checked them if I could hear anything during the time and sometimes there was no noise between the interval [t_min, t_max]",
    "1104378": "Unless t_min == t_max, you should hear audio. Since it comes from a continuous 60-second audio clip, there will always be sound (even if it's just background noise). Just a thought!",
    "1104382": "I do hear audio, but I dont hear the species which should be making noise during this interval.",
    "1109674": "Another augmentation method for audio is reverberation"
  },
  "source": "meta"
}