{
  "id": 91431,
  "title": "Paper: Deep Learning for Audio Signal Processing",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/91431",
  "author_name": "",
  "post_date": "2019-05-04T15:37:45.004216600Z",
  "votes": 41,
  "comment_count": 9,
  "views": 0,
  "content": "<p><a href=\"https://arxiv.org/pdf/1905.00078.pdf\">https://arxiv.org/pdf/1905.00078.pdf</a>\nThis is also from twitter, very informative paper that provides a review of the SOTA deep learning techniques for audio signal processing. My notes for this competition.</p>\n\n<ul>\n<li>B. Audio Features, <em>'... However, due to the physics of sound production, there are additional correlations for frequencies that are multiples of the same base frequency (harmonics). To allow a spatially local model (e.g., a CNN) to take these into account, a third dimension can be added that directly yields the magnitudes of the harmonic series [14], [15].'</em> -- Very interesting topic.</li>\n<li>C. Models, various models are discussed. Useful, but a lot to do...</li>\n<li><em>'f) Phase modeling: In the calculation of the log-mel spectrum, the magnitude spectrum is used but the phase spectrum is lost. While this may be desired for analysis, synthesis requires plausible phases. ...'</em> -- Phase is one of the things I'm trying to make use of, it would be of some help in my stacking ensemble.</li>\n<li>D. Data, data generation and augmentation are discussed.</li>\n<li>... And many more topics.</li>\n</ul>",
  "messages": [
    {
      "id": "527106",
      "postDate": "05/04/2019 15:37:45",
      "content": "<p><a href=\"https://arxiv.org/pdf/1905.00078.pdf\">https://arxiv.org/pdf/1905.00078.pdf</a>\nThis is also from twitter, very informative paper that provides a review of the SOTA deep learning techniques for audio signal processing. My notes for this competition.</p>\n\n<ul>\n<li>B. Audio Features, <em>'... However, due to the physics of sound production, there are additional correlations for frequencies that are multiples of the same base frequency (harmonics). To allow a spatially local model (e.g., a CNN) to take these into account, a third dimension can be added that directly yields the magnitudes of the harmonic series [14], [15].'</em> -- Very interesting topic.</li>\n<li>C. Models, various models are discussed. Useful, but a lot to do...</li>\n<li><em>'f) Phase modeling: In the calculation of the log-mel spectrum, the magnitude spectrum is used but the phase spectrum is lost. While this may be desired for analysis, synthesis requires plausible phases. ...'</em> -- Phase is one of the things I'm trying to make use of, it would be of some help in my stacking ensemble.</li>\n<li>D. Data, data generation and augmentation are discussed.</li>\n<li>... And many more topics.</li>\n</ul>",
      "rawMarkdown": "https://arxiv.org/pdf/1905.00078.pdf\nThis is also from twitter, very informative paper that provides a review of the SOTA deep learning techniques for audio signal processing. My notes for this competition.\n\n- B. Audio Features, _'... However, due to the physics of sound production, there are additional correlations for frequencies that are multiples of the same base frequency (harmonics). To allow a spatially local model (e.g., a CNN) to take these into account, a third dimension can be added that directly yields the magnitudes of the harmonic series [14], [15].'_ -- Very interesting topic.\n- C. Models, various models are discussed. Useful, but a lot to do...\n- _'f) Phase modeling: In the calculation of the log-mel spectrum, the magnitude spectrum is used but the phase spectrum is lost. While this may be desired for analysis, synthesis requires plausible phases. ...'_ -- Phase is one of the things I'm trying to make use of, it would be of some help in my stacking ensemble.\n- D. Data, data generation and augmentation are discussed.\n- ... And many more topics.",
      "votes": null
    },
    {
      "id": "527368",
      "postDate": "05/05/2019 08:34:38",
      "content": "<p>Very useful! Thanks for sharing!</p>",
      "rawMarkdown": "Very useful! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "527504",
      "postDate": "05/05/2019 17:14:57",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "527740",
      "postDate": "05/06/2019 07:56:12",
      "content": "<p>Thank you for posting this </p>",
      "rawMarkdown": "Thank you for posting this",
      "votes": null
    },
    {
      "id": "528346",
      "postDate": "05/07/2019 14:42:31",
      "content": "<p>Thanks a lot for sharing your insights!</p>",
      "rawMarkdown": "Thanks a lot for sharing your insights!",
      "votes": null
    },
    {
      "id": "528896",
      "postDate": "05/08/2019 21:46:49",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "529890",
      "postDate": "05/11/2019 04:40:30",
      "content": "<p>Great! Thanks for sharing!</p>",
      "rawMarkdown": "Great! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "530409",
      "postDate": "05/12/2019 18:21:22",
      "content": "<p>Thank you kindly for sharing! ♥\nThis paper is really insightful.</p>",
      "rawMarkdown": "Thank you kindly for sharing! ♥\nThis paper is really insightful.",
      "votes": null
    },
    {
      "id": "543804",
      "postDate": "06/04/2019 20:57:08",
      "content": "<p>I tried Phase modeling. I generated phase waves with Matplotlib, but the images generated are too similar for a model to perform a satisfying classification on these images. I achieved a mere 0.15 lwlrap on the dataset generated. I don't think it is relevant to include it in a stacked ensemble. </p>",
      "rawMarkdown": "I tried Phase modeling. I generated phase waves with Matplotlib, but the images generated are too similar for a model to perform a satisfying classification on these images. I achieved a mere 0.15 lwlrap on the dataset generated. I don't think it is relevant to include it in a stacked ensemble.",
      "votes": null
    },
    {
      "id": "544038",
      "postDate": "06/05/2019 04:02:20",
      "content": "<p>Hi I agree with your observation, it doesn't give us huge jump. Currently I'm using 3 channels - amplitude, phase and one more derived from these values - as basic features for multi label CNN models. I hope somebody else finds ways to make better use of phase!</p>",
      "rawMarkdown": "Hi I agree with your observation, it doesn't give us huge jump. Currently I'm using 3 channels - amplitude, phase and one more derived from these values - as basic features for multi label CNN models. I hope somebody else finds ways to make better use of phase!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 527368,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "05/05/2019 08:34:38",
      "content": "<p>Very useful! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 527504,
      "author_name": "wuzuping",
      "author_url": "",
      "post_date": "05/05/2019 17:14:57",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 527740,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "05/06/2019 07:56:12",
      "content": "<p>Thank you for posting this </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528346,
      "author_name": "praveern",
      "author_url": "",
      "post_date": "05/07/2019 14:42:31",
      "content": "<p>Thanks a lot for sharing your insights!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528896,
      "author_name": "yussiroz",
      "author_url": "",
      "post_date": "05/08/2019 21:46:49",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 529890,
      "author_name": "thiventura",
      "author_url": "",
      "post_date": "05/11/2019 04:40:30",
      "content": "<p>Great! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 530409,
      "author_name": "varnez",
      "author_url": "",
      "post_date": "05/12/2019 18:21:22",
      "content": "<p>Thank you kindly for sharing! ♥\nThis paper is really insightful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543804,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/04/2019 20:57:08",
      "content": "<p>I tried Phase modeling. I generated phase waves with Matplotlib, but the images generated are too similar for a model to perform a satisfying classification on these images. I achieved a mere 0.15 lwlrap on the dataset generated. I don't think it is relevant to include it in a stacked ensemble. </p>",
      "votes": null,
      "replies": [
        {
          "id": 544038,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/05/2019 04:02:20",
          "content": "<p>Hi I agree with your observation, it doesn't give us huge jump. Currently I'm using 3 channels - amplitude, phase and one more derived from these values - as basic features for multi label CNN models. I hope somebody else finds ways to make better use of phase!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "527106": "https://arxiv.org/pdf/1905.00078.pdf\nThis is also from twitter, very informative paper that provides a review of the SOTA deep learning techniques for audio signal processing. My notes for this competition.\n\n- B. Audio Features, _'... However, due to the physics of sound production, there are additional correlations for frequencies that are multiples of the same base frequency (harmonics). To allow a spatially local model (e.g., a CNN) to take these into account, a third dimension can be added that directly yields the magnitudes of the harmonic series [14], [15].'_ -- Very interesting topic.\n- C. Models, various models are discussed. Useful, but a lot to do...\n- _'f) Phase modeling: In the calculation of the log-mel spectrum, the magnitude spectrum is used but the phase spectrum is lost. While this may be desired for analysis, synthesis requires plausible phases. ...'_ -- Phase is one of the things I'm trying to make use of, it would be of some help in my stacking ensemble.\n- D. Data, data generation and augmentation are discussed.\n- ... And many more topics.",
    "527368": "Very useful! Thanks for sharing!",
    "527504": "Thanks for sharing!",
    "527740": "Thank you for posting this",
    "528346": "Thanks a lot for sharing your insights!",
    "528896": "",
    "529890": "Great! Thanks for sharing!",
    "530409": "Thank you kindly for sharing! ♥\nThis paper is really insightful.",
    "543804": "I tried Phase modeling. I generated phase waves with Matplotlib, but the images generated are too similar for a model to perform a satisfying classification on these images. I achieved a mere 0.15 lwlrap on the dataset generated. I don't think it is relevant to include it in a stacked ensemble.",
    "544038": "Hi I agree with your observation, it doesn't give us huge jump. Currently I'm using 3 channels - amplitude, phase and one more derived from these values - as basic features for multi label CNN models. I hope somebody else finds ways to make better use of phase!"
  },
  "source": "meta"
}