{
  "id": 90786,
  "title": "Reverse Spectogram",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/90786",
  "author_name": "",
  "post_date": "2019-04-27T08:12:11.760889400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>One of the topics I'd like to explore is the validity of data augmentation (done in spectogram space). If I do random erasing as is suggested <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/90464#latest-523788\">here</a>, what would happen to the actual audio for instance.</p>\n\n<p>So from <a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\">this kernel</a> it can be seen how to get the spectograms using librosa. Is there a way to go from spectogram, back to the audio file?</p>",
  "messages": [
    {
      "id": "523853",
      "postDate": "04/27/2019 08:12:11",
      "content": "<p>One of the topics I'd like to explore is the validity of data augmentation (done in spectogram space). If I do random erasing as is suggested <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/90464#latest-523788\">here</a>, what would happen to the actual audio for instance.</p>\n\n<p>So from <a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\">this kernel</a> it can be seen how to get the spectograms using librosa. Is there a way to go from spectogram, back to the audio file?</p>",
      "rawMarkdown": "One of the topics I'd like to explore is the validity of data augmentation (done in spectogram space). If I do random erasing as is suggested [here](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/90464#latest-523788), what would happen to the actual audio for instance.\n\nSo from [this kernel](https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai) it can be seen how to get the spectograms using librosa. Is there a way to go from spectogram, back to the audio file?",
      "votes": null
    },
    {
      "id": "523866",
      "postDate": "04/27/2019 08:49:30",
      "content": "<p>Going from a magnitude spectrogram to raw audio is tricky and currently it is an active area of research. There is a simple algorithm called Griffin-Lim which iteratively estimates the phase from the magnitude. However, the sound quality produced by Griffin-Lim is, in the best case, mediocre. That's why many neural vocoders have been introduced:</p>\n\n<p><a href=\"https://arxiv.org/abs/1609.03499\">https://arxiv.org/abs/1609.03499</a>\n<a href=\"https://arxiv.org/abs/1811.00002\">https://arxiv.org/abs/1811.00002</a>\n<a href=\"https://arxiv.org/abs/1811.02155\">https://arxiv.org/abs/1811.02155</a></p>\n\n<p>However, if you have access to the phase, which is typically discarded after spectrogram calculation, you can simply reconstruct the waveform using the inverse Fourier transform.</p>",
      "rawMarkdown": "Going from a magnitude spectrogram to raw audio is tricky and currently it is an active area of research. There is a simple algorithm called Griffin-Lim which iteratively estimates the phase from the magnitude. However, the sound quality produced by Griffin-Lim is, in the best case, mediocre. That's why many neural vocoders have been introduced:\n\nhttps://arxiv.org/abs/1609.03499\nhttps://arxiv.org/abs/1811.00002\nhttps://arxiv.org/abs/1811.02155\n\nHowever, if you have access to the phase, which is typically discarded after spectrogram calculation, you can simply reconstruct the waveform using the inverse Fourier transform.",
      "votes": null
    },
    {
      "id": "523892",
      "postDate": "04/27/2019 09:58:18",
      "content": "<p>You can also try to play with some audio editing software to gain some intuition about the spectograms. When you do random erasing with vertical bands it's just introducing silence in that region. For horizontal bands is filtering out some frequencies. My intuition is that it should help in particular on multi-label noisy cases where multiple sounds coexist and the horizontal erasing can mask some. </p>",
      "rawMarkdown": "You can also try to play with some audio editing software to gain some intuition about the spectograms. When you do random erasing with vertical bands it's just introducing silence in that region. For horizontal bands is filtering out some frequencies. My intuition is that it should help in particular on multi-label noisy cases where multiple sounds coexist and the horizontal erasing can mask some.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 523866,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "04/27/2019 08:49:30",
      "content": "<p>Going from a magnitude spectrogram to raw audio is tricky and currently it is an active area of research. There is a simple algorithm called Griffin-Lim which iteratively estimates the phase from the magnitude. However, the sound quality produced by Griffin-Lim is, in the best case, mediocre. That's why many neural vocoders have been introduced:</p>\n\n<p><a href=\"https://arxiv.org/abs/1609.03499\">https://arxiv.org/abs/1609.03499</a>\n<a href=\"https://arxiv.org/abs/1811.00002\">https://arxiv.org/abs/1811.00002</a>\n<a href=\"https://arxiv.org/abs/1811.02155\">https://arxiv.org/abs/1811.02155</a></p>\n\n<p>However, if you have access to the phase, which is typically discarded after spectrogram calculation, you can simply reconstruct the waveform using the inverse Fourier transform.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 523892,
      "author_name": "mnpinto",
      "author_url": "",
      "post_date": "04/27/2019 09:58:18",
      "content": "<p>You can also try to play with some audio editing software to gain some intuition about the spectograms. When you do random erasing with vertical bands it's just introducing silence in that region. For horizontal bands is filtering out some frequencies. My intuition is that it should help in particular on multi-label noisy cases where multiple sounds coexist and the horizontal erasing can mask some. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "523853": "One of the topics I'd like to explore is the validity of data augmentation (done in spectogram space). If I do random erasing as is suggested [here](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/90464#latest-523788), what would happen to the actual audio for instance.\n\nSo from [this kernel](https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai) it can be seen how to get the spectograms using librosa. Is there a way to go from spectogram, back to the audio file?",
    "523866": "Going from a magnitude spectrogram to raw audio is tricky and currently it is an active area of research. There is a simple algorithm called Griffin-Lim which iteratively estimates the phase from the magnitude. However, the sound quality produced by Griffin-Lim is, in the best case, mediocre. That's why many neural vocoders have been introduced:\n\nhttps://arxiv.org/abs/1609.03499\nhttps://arxiv.org/abs/1811.00002\nhttps://arxiv.org/abs/1811.02155\n\nHowever, if you have access to the phase, which is typically discarded after spectrogram calculation, you can simply reconstruct the waveform using the inverse Fourier transform.",
    "523892": "You can also try to play with some audio editing software to gain some intuition about the spectograms. When you do random erasing with vertical bands it's just introducing silence in that region. For horizontal bands is filtering out some frequencies. My intuition is that it should help in particular on multi-label noisy cases where multiple sounds coexist and the horizontal erasing can mask some."
  },
  "source": "meta"
}