{
  "id": 491534,
  "title": "Pseudolabeling Youtube Speech Data",
  "url": "/competitions/ben10/discussion/491534",
  "author_name": "Tahsin",
  "post_date": "2024-04-06T07:51:21.286000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The dataset we are working with is comparatively small. We can pseudolabel external data to increase the dataset size.</p>\n<p>In the past Bengali ASR competition, Youtube data was pseudolabeled to improve the performance of models.</p>\n<p>Pseudolabeling involves the following three steps.</p>\n<p>Step 1, we can create a list of Youtube channels dialect speech and download audios from those channels.<br>\n<a href=\"https://www.kaggle.com/code/reasat/pseudolabeling-step-1-download-speech-audio\" target=\"_blank\">https://www.kaggle.com/code/reasat/pseudolabeling-step-1-download-speech-audio</a></p>\n<p>Step 2, run a voice activity detector model to extract chunks of audio containing speech.<br>\n<a href=\"https://www.kaggle.com/code/reasat/pseudolabeling-step-2-create-speech-chunks\" target=\"_blank\">https://www.kaggle.com/code/reasat/pseudolabeling-step-2-create-speech-chunks</a></p>\n<p>Step 3,  use an existing ASR model to create pseudo transcription for the audio chunks.<br>\n<a href=\"https://www.kaggle.com/code/reasat/pseudolabeling-step-3-infer\" target=\"_blank\">https://www.kaggle.com/code/reasat/pseudolabeling-step-3-infer</a></p>\n<p>Rather then individual people trying to cover random subsets of Youtube channels, we can coordinate among ourselves to create a large pseudolabeled spontaneous speech dataset.</p>\n<p>Let me know if you would like to contribute in making a list of channels and taking the lead in making pseudolabels. Here's a csv file to coordinate the pseudo labeling initiative <a href=\"https://docs.google.com/spreadsheets/d/1M3JPvXFF85Qska6Ob5UGj1f-OwgW4nry1d2AoQnV6mU/edit?usp=sharing\" target=\"_blank\">csv link</a>, I'll update it as I find potential channels.</p>",
  "messages": [
    {
      "id": 2738230,
      "postDate": "2024-04-06T07:51:21.287Z",
      "content": "<p>The dataset we are working with is comparatively small. We can pseudolabel external data to increase the dataset size.</p>\n<p>In the past Bengali ASR competition, Youtube data was pseudolabeled to improve the performance of models.</p>\n<p>Pseudolabeling involves the following three steps.</p>\n<p>Step 1, we can create a list of Youtube channels dialect speech and download audios from those channels.<br>\n<a href=\"https://www.kaggle.com/code/reasat/pseudolabeling-step-1-download-speech-audio\" target=\"_blank\">https://www.kaggle.com/code/reasat/pseudolabeling-step-1-download-speech-audio</a></p>\n<p>Step 2, run a voice activity detector model to extract chunks of audio containing speech.<br>\n<a href=\"https://www.kaggle.com/code/reasat/pseudolabeling-step-2-create-speech-chunks\" target=\"_blank\">https://www.kaggle.com/code/reasat/pseudolabeling-step-2-create-speech-chunks</a></p>\n<p>Step 3,  use an existing ASR model to create pseudo transcription for the audio chunks.<br>\n<a href=\"https://www.kaggle.com/code/reasat/pseudolabeling-step-3-infer\" target=\"_blank\">https://www.kaggle.com/code/reasat/pseudolabeling-step-3-infer</a></p>\n<p>Rather then individual people trying to cover random subsets of Youtube channels, we can coordinate among ourselves to create a large pseudolabeled spontaneous speech dataset.</p>\n<p>Let me know if you would like to contribute in making a list of channels and taking the lead in making pseudolabels. Here's a csv file to coordinate the pseudo labeling initiative <a href=\"https://docs.google.com/spreadsheets/d/1M3JPvXFF85Qska6Ob5UGj1f-OwgW4nry1d2AoQnV6mU/edit?usp=sharing\" target=\"_blank\">csv link</a>, I'll update it as I find potential channels.</p>",
      "rawMarkdown": "The dataset we are working with is comparatively small. We can pseudolabel external data to increase the dataset size.\n\nIn the past Bengali ASR competition, Youtube data was pseudolabeled to improve the performance of models.\n\nPseudolabeling involves the following three steps.\n\nStep 1, we can create a list of Youtube channels dialect speech and download audios from those channels.\nhttps://www.kaggle.com/code/reasat/pseudolabeling-step-1-download-speech-audio\n\nStep 2, run a voice activity detector model to extract chunks of audio containing speech.\nhttps://www.kaggle.com/code/reasat/pseudolabeling-step-2-create-speech-chunks\n\nStep 3,  use an existing ASR model to create pseudo transcription for the audio chunks.\nhttps://www.kaggle.com/code/reasat/pseudolabeling-step-3-infer\n\nRather then individual people trying to cover random subsets of Youtube channels, we can coordinate among ourselves to create a large pseudolabeled spontaneous speech dataset.\n\nLet me know if you would like to contribute in making a list of channels and taking the lead in making pseudolabels. Here's a csv file to coordinate the pseudo labeling initiative [csv link](https://docs.google.com/spreadsheets/d/1M3JPvXFF85Qska6Ob5UGj1f-OwgW4nry1d2AoQnV6mU/edit?usp=sharing), I'll update it as I find potential channels.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2738230": "The dataset we are working with is comparatively small. We can pseudolabel external data to increase the dataset size.\n\nIn the past Bengali ASR competition, Youtube data was pseudolabeled to improve the performance of models.\n\nPseudolabeling involves the following three steps.\n\nStep 1, we can create a list of Youtube channels dialect speech and download audios from those channels.\nhttps://www.kaggle.com/code/reasat/pseudolabeling-step-1-download-speech-audio\n\nStep 2, run a voice activity detector model to extract chunks of audio containing speech.\nhttps://www.kaggle.com/code/reasat/pseudolabeling-step-2-create-speech-chunks\n\nStep 3,  use an existing ASR model to create pseudo transcription for the audio chunks.\nhttps://www.kaggle.com/code/reasat/pseudolabeling-step-3-infer\n\nRather then individual people trying to cover random subsets of Youtube channels, we can coordinate among ourselves to create a large pseudolabeled spontaneous speech dataset.\n\nLet me know if you would like to contribute in making a list of channels and taking the lead in making pseudolabels. Here's a csv file to coordinate the pseudo labeling initiative [csv link](https://docs.google.com/spreadsheets/d/1M3JPvXFF85Qska6Ob5UGj1f-OwgW4nry1d2AoQnV6mU/edit?usp=sharing), I'll update it as I find potential channels."
  }
}