{
  "id": 440964,
  "title": "Speech Enhancement",
  "url": "/competitions/bengaliai-speech/discussion/440964",
  "author_name": "Roy Wei",
  "post_date": "2023-09-17T02:00:22.173000",
  "votes": 1,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I have noticed that <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/434179\" target=\"_blank\">Facebook_denoiser</a> is not allowed to be used in the competition because of the license. What are some alternative solutions? </p>\n<p>I've experimented with several SpeechBrain models on Huggingface, but upon listening, the results were not particularly impressive.</p>\n<p>Should I train it from scratch or is their other free-to-use state of art model for usage?</p>\n<p>Thanks a lot to anyone who would like to share their opinions</p>",
  "messages": [
    {
      "id": 2442390,
      "postDate": "2023-09-17T02:00:22.173Z",
      "content": "<p>I have noticed that <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/434179\" target=\"_blank\">Facebook_denoiser</a> is not allowed to be used in the competition because of the license. What are some alternative solutions? </p>\n<p>I've experimented with several SpeechBrain models on Huggingface, but upon listening, the results were not particularly impressive.</p>\n<p>Should I train it from scratch or is their other free-to-use state of art model for usage?</p>\n<p>Thanks a lot to anyone who would like to share their opinions</p>",
      "rawMarkdown": "I have noticed that [Facebook_denoiser](https://www.kaggle.com/competitions/bengaliai-speech/discussion/434179) is not allowed to be used in the competition because of the license. What are some alternative solutions? \n\nI've experimented with several SpeechBrain models on Huggingface, but upon listening, the results were not particularly impressive.\n\nShould I train it from scratch or is their other free-to-use state of art model for usage?\n\nThanks a lot to anyone who would like to share their opinions",
      "votes": 1
    },
    {
      "id": 2442548,
      "postDate": "2023-09-17T05:36:03.353Z",
      "content": "<p>Try this one : <a href=\"https://github.com/NVIDIA/CleanUNet\" target=\"_blank\">https://github.com/NVIDIA/CleanUNet</a></p>",
      "rawMarkdown": "Try this one : [https://github.com/NVIDIA/CleanUNet](https://github.com/NVIDIA/CleanUNet)",
      "replies": [
        {
          "id": 2443079,
          "postDate": "2023-09-17T14:09:39.633Z",
          "content": "<p>Hello, thank you so much for your suggestion! I've given several denoising models a try, but unfortunately, they didn't quite hit the mark on the LB. It seems that denoising might be causing some harm to the features, which complicates the model's prediction task. By any chance, have you had the opportunity to test this model? I'm curious if it has yielded better results for you?</p>",
          "rawMarkdown": "Hello, thank you so much for your suggestion! I've given several denoising models a try, but unfortunately, they didn't quite hit the mark on the LB. It seems that denoising might be causing some harm to the features, which complicates the model's prediction task. By any chance, have you had the opportunity to test this model? I'm curious if it has yielded better results for you?",
          "replies": [
            {
              "id": 2443107,
              "postDate": "2023-09-17T14:22:17.793Z",
              "content": "<p>In my previous experiment, it improved my score by 0.02</p>",
              "rawMarkdown": "In my previous experiment, it improved my score by 0.02",
              "votes": 1
            },
            {
              "id": 2443156,
              "postDate": "2023-09-17T14:55:43.330Z",
              "content": "<p>I have tried this <a href=\"https://huggingface.co/speechbrain/metricgan-plus-voicebank\" target=\"_blank\">denoiser</a>, it did denoise the background noise, but also muffle human voice as well. </p>",
              "rawMarkdown": "I have tried this [denoiser](https://huggingface.co/speechbrain/metricgan-plus-voicebank), it did denoise the background noise, but also muffle human voice as well. ",
              "votes": 1
            },
            {
              "id": 2447192,
              "postDate": "2023-09-19T22:24:23.497Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2447193,
              "postDate": "2023-09-19T22:25:13.747Z",
              "content": "<p>I have implemented the CeanUnet, but the notebook timeout. Do you mind giving me some hint of what to do. I have used p100 and dns large full pretrained checkpoint</p>",
              "rawMarkdown": "I have implemented the CeanUnet, but the notebook timeout. Do you mind giving me some hint of what to do. I have used p100 and dns large full pretrained checkpoint"
            },
            {
              "id": 2447512,
              "postDate": "2023-09-20T06:01:19.380Z",
              "content": "<p>Me and my teammate tried to ways to do it: </p>\n<p>The first way was to denoise all the audios, save them then reload them. But that ran out of time.</p>\n<p>The second way was to denoise while inference, we denoise before every batch infer. That works, the notebook took about 8 hours to run (The main inference took about 5 hours). Although that have not yeild a better score for us yet:(</p>\n<p>The model loaded in the GPU occupied some space, but it was still faster. You should try this if you used the first way, which required to save and load the audio file.</p>",
              "rawMarkdown": "Me and my teammate tried to ways to do it: \n\nThe first way was to denoise all the audios, save them then reload them. But that ran out of time.\n\n\nThe second way was to denoise while inference, we denoise before every batch infer. That works, the notebook took about 8 hours to run (The main inference took about 5 hours). Although that have not yeild a better score for us yet:(\n\nThe model loaded in the GPU occupied some space, but it was still faster. You should try this if you used the first way, which required to save and load the audio file."
            },
            {
              "id": 2447951,
              "postDate": "2023-09-20T10:41:17.870Z",
              "content": "<p>I'm sorry but I don't quite get your kind advice. My way is just to denoise with the CleanUnet and store the denoised audios int the \"working\" directory, then load them  into the model for inference, yet it sadly timeouts. Is that what you mean in the first scenario?</p>",
              "rawMarkdown": "I'm sorry but I don't quite get your kind advice. My way is just to denoise with the CleanUnet and store the denoised audios int the \"working\" directory, then load them  into the model for inference, yet it sadly timeouts. Is that what you mean in the first scenario?"
            },
            {
              "id": 2447952,
              "postDate": "2023-09-20T10:44:01.440Z",
              "content": "<p>Teammate here, tried to implement CleanUNet, only to find a not better score. I wonder if denoising does some harm to inference, like reducing important features? Maybe it's the fault of using both denoising and punctuation restoring model?</p>",
              "rawMarkdown": "Teammate here, tried to implement CleanUNet, only to find a not better score. I wonder if denoising does some harm to inference, like reducing important features? Maybe it's the fault of using both denoising and punctuation restoring model?",
              "votes": 1
            },
            {
              "id": 2447955,
              "postDate": "2023-09-20T10:46:13.437Z",
              "content": "<p>Yes, he meant that exactly. Later, we'd just extracted a method from the original nvidia project, and it worked within 9h well, but the perf didn't get better.</p>",
              "rawMarkdown": "Yes, he meant that exactly. Later, we'd just extracted a method from the original nvidia project, and it worked within 9h well, but the perf didn't get better."
            },
            {
              "id": 2447958,
              "postDate": "2023-09-20T10:48:33.327Z",
              "content": "<p>In detail, 0.400-&gt;0.426 after implementation of CleanUNet. Wondering if there were some wrong settings in my notebook…</p>",
              "rawMarkdown": "In detail, 0.400->0.426 after implementation of CleanUNet. Wondering if there were some wrong settings in my notebook..."
            },
            {
              "id": 2447963,
              "postDate": "2023-09-20T10:53:13.780Z",
              "content": "<p>Feel sorry for you guys… I think I will then try FAIR denoiser first, since I am not aiming at the prize😅. I will post my observations here, if I am able to implement Facebook's model</p>",
              "rawMarkdown": "Feel sorry for you guys... I think I will then try FAIR denoiser first, since I am not aiming at the prize😅. I will post my observations here, if I am able to implement Facebook's model",
              "votes": 1
            },
            {
              "id": 2447993,
              "postDate": "2023-09-20T11:13:33.507Z",
              "content": "<p>Is the input audio being resampled to a 16,000 Hz sampling rate and restricted to only one channel?</p>",
              "rawMarkdown": "Is the input audio being resampled to a 16,000 Hz sampling rate and restricted to only one channel?"
            },
            {
              "id": 2448007,
              "postDate": "2023-09-20T11:22:10.987Z",
              "content": "<p><code>mono=True</code> in <code>load</code> function from <code>librosa</code>, for example?</p>",
              "rawMarkdown": "`mono=True` in `load` function from `librosa`, for example?"
            },
            {
              "id": 2448380,
              "postDate": "2023-09-20T14:53:08.533Z",
              "content": "<p>Yes, I faced this issue for the first time. After fixing it, my score improved from 0.428 to 0.426. Regarding Librosa's resampling method, I suggest using res_type='fft'. After this experiment, I never attempted to remove it, so I'm not sure if it would still be beneficial when the speech recognition model becomes more robust.</p>",
              "rawMarkdown": "Yes, I faced this issue for the first time. After fixing it, my score improved from 0.428 to 0.426. Regarding Librosa's resampling method, I suggest using res_type='fft'. After this experiment, I never attempted to remove it, so I'm not sure if it would still be beneficial when the speech recognition model becomes more robust."
            },
            {
              "id": 2449238,
              "postDate": "2023-09-21T04:51:22.540Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2449731,
              "postDate": "2023-09-21T11:46:50.713Z",
              "content": "<p>resampling rate is meaning <code>sr</code> param in <code>librosa.load</code>? or another setting…</p>",
              "rawMarkdown": "resampling rate is meaning `sr` param in `librosa.load`? or another setting..."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2442548,
      "author_name": "Thananchai Kongthaworn",
      "author_url": "",
      "post_date": "2023-09-17T05:36:03.353000",
      "content": "<p>Try this one : <a href=\"https://github.com/NVIDIA/CleanUNet\" target=\"_blank\">https://github.com/NVIDIA/CleanUNet</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2443079,
          "author_name": "Mutian Hong",
          "author_url": "",
          "post_date": "2023-09-17T14:09:39.633000",
          "content": "<p>Hello, thank you so much for your suggestion! I've given several denoising models a try, but unfortunately, they didn't quite hit the mark on the LB. It seems that denoising might be causing some harm to the features, which complicates the model's prediction task. By any chance, have you had the opportunity to test this model? I'm curious if it has yielded better results for you?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2443107,
              "author_name": "Thananchai Kongthaworn",
              "author_url": "",
              "post_date": "2023-09-17T14:22:17.793000",
              "content": "<p>In my previous experiment, it improved my score by 0.02</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2443156,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2023-09-17T14:55:43.330000",
              "content": "<p>I have tried this <a href=\"https://huggingface.co/speechbrain/metricgan-plus-voicebank\" target=\"_blank\">denoiser</a>, it did denoise the background noise, but also muffle human voice as well. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2447192,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-19T22:24:23.497000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2447193,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2023-09-19T22:25:13.747000",
              "content": "<p>I have implemented the CeanUnet, but the notebook timeout. Do you mind giving me some hint of what to do. I have used p100 and dns large full pretrained checkpoint</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2447512,
              "author_name": "Mutian Hong",
              "author_url": "",
              "post_date": "2023-09-20T06:01:19.380000",
              "content": "<p>Me and my teammate tried to ways to do it: </p>\n<p>The first way was to denoise all the audios, save them then reload them. But that ran out of time.</p>\n<p>The second way was to denoise while inference, we denoise before every batch infer. That works, the notebook took about 8 hours to run (The main inference took about 5 hours). Although that have not yeild a better score for us yet:(</p>\n<p>The model loaded in the GPU occupied some space, but it was still faster. You should try this if you used the first way, which required to save and load the audio file.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2447951,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2023-09-20T10:41:17.870000",
              "content": "<p>I'm sorry but I don't quite get your kind advice. My way is just to denoise with the CleanUnet and store the denoised audios int the \"working\" directory, then load them  into the model for inference, yet it sadly timeouts. Is that what you mean in the first scenario?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2447952,
              "author_name": "iiiiitsu",
              "author_url": "",
              "post_date": "2023-09-20T10:44:01.440000",
              "content": "<p>Teammate here, tried to implement CleanUNet, only to find a not better score. I wonder if denoising does some harm to inference, like reducing important features? Maybe it's the fault of using both denoising and punctuation restoring model?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2447955,
              "author_name": "iiiiitsu",
              "author_url": "",
              "post_date": "2023-09-20T10:46:13.437000",
              "content": "<p>Yes, he meant that exactly. Later, we'd just extracted a method from the original nvidia project, and it worked within 9h well, but the perf didn't get better.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2447958,
              "author_name": "iiiiitsu",
              "author_url": "",
              "post_date": "2023-09-20T10:48:33.327000",
              "content": "<p>In detail, 0.400-&gt;0.426 after implementation of CleanUNet. Wondering if there were some wrong settings in my notebook…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2447963,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2023-09-20T10:53:13.780000",
              "content": "<p>Feel sorry for you guys… I think I will then try FAIR denoiser first, since I am not aiming at the prize😅. I will post my observations here, if I am able to implement Facebook's model</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2447993,
              "author_name": "Thananchai Kongthaworn",
              "author_url": "",
              "post_date": "2023-09-20T11:13:33.507000",
              "content": "<p>Is the input audio being resampled to a 16,000 Hz sampling rate and restricted to only one channel?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2448007,
              "author_name": "iiiiitsu",
              "author_url": "",
              "post_date": "2023-09-20T11:22:10.987000",
              "content": "<p><code>mono=True</code> in <code>load</code> function from <code>librosa</code>, for example?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2448380,
              "author_name": "Thananchai Kongthaworn",
              "author_url": "",
              "post_date": "2023-09-20T14:53:08.533000",
              "content": "<p>Yes, I faced this issue for the first time. After fixing it, my score improved from 0.428 to 0.426. Regarding Librosa's resampling method, I suggest using res_type='fft'. After this experiment, I never attempted to remove it, so I'm not sure if it would still be beneficial when the speech recognition model becomes more robust.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2449238,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-21T04:51:22.540000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2449731,
              "author_name": "iiiiitsu",
              "author_url": "",
              "post_date": "2023-09-21T11:46:50.713000",
              "content": "<p>resampling rate is meaning <code>sr</code> param in <code>librosa.load</code>? or another setting…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2442390": "I have noticed that [Facebook_denoiser](https://www.kaggle.com/competitions/bengaliai-speech/discussion/434179) is not allowed to be used in the competition because of the license. What are some alternative solutions? \n\nI've experimented with several SpeechBrain models on Huggingface, but upon listening, the results were not particularly impressive.\n\nShould I train it from scratch or is their other free-to-use state of art model for usage?\n\nThanks a lot to anyone who would like to share their opinions",
    "2442548": "Try this one : [https://github.com/NVIDIA/CleanUNet](https://github.com/NVIDIA/CleanUNet)"
  }
}