{
  "id": 427218,
  "title": "Introducing AI voice remover and AI voice changer",
  "url": "/competitions/bengaliai-speech/discussion/427218",
  "author_name": "nadare",
  "post_date": "2023-07-26T23:27:25.854000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello Kaggler</p>\n<p>I will introduce free software that seems to be useful for speech recognition.</p>\n<h1>AI voice remover</h1>\n<p>The voice of the training data for this competition contains some voices other than human voices.<br>\nFor this, you can separate the environmental sound and the human voice by using the ultimate vocal remover.<br>\n<a href=\"https://ultimatevocalremover.com/\" target=\"_blank\">https://ultimatevocalremover.com/</a></p>\n<h1>AI voice changer</h1>\n<p>Next is the introduction of voice changer.</p>\n<h2>RVC-webUI</h2>\n<p>RVC has become a hot topic recently, but this is an operation over the webui.<br>\nddPn08's RVC-fork can convert audio using webAPI while running RVC-webUI locally. If you are familiar with RVC, please try it.<br>\n<a href=\"https://github.com/ddPn08/rvc-webui/tree/main\" target=\"_blank\">https://github.com/ddPn08/rvc-webui/tree/main</a></p>\n<h2>VoRAS</h2>\n<p>Most of RVC converts to one speaker with one model, but in my library called voras, which was created by remodeling RVC, speech conversion to more than 2000 people trained using LibriTTS-R can be performed at once. It is possible by simply switching the input id in one model. Use it for data augmentation.<br>\n<a href=\"https://github.com/nadare881/voras-webui-beta/tree/main\" target=\"_blank\">https://github.com/nadare881/voras-webui-beta/tree/main</a></p>\n<p>voras_pretrain_libritts_r.pth is a pretrain model trained with over 2000 speakers.<br>\n<a href=\"https://huggingface.co/datasets/nadare/voras/tree/main\" target=\"_blank\">https://huggingface.co/datasets/nadare/voras/tree/main</a></p>\n<p>This is a demonstration of speech conversion using a model trained by five Japanese free sound sources, but it is also possible to convert speech through webAPI.<br>\n<a href=\"https://twitter.com/i/status/1682371135718199296\" target=\"_blank\">https://twitter.com/i/status/1682371135718199296</a></p>\n<p>Voras plans to leave it without any fixes other than bugfixes at least until the end of the competition. If you have any questions about how to use voras ask me in this thread. (I don't know if I see them often)</p>",
  "messages": [
    {
      "id": 2360698,
      "postDate": "2023-07-26T23:27:25.853Z",
      "content": "<p>Hello Kaggler</p>\n<p>I will introduce free software that seems to be useful for speech recognition.</p>\n<h1>AI voice remover</h1>\n<p>The voice of the training data for this competition contains some voices other than human voices.<br>\nFor this, you can separate the environmental sound and the human voice by using the ultimate vocal remover.<br>\n<a href=\"https://ultimatevocalremover.com/\" target=\"_blank\">https://ultimatevocalremover.com/</a></p>\n<h1>AI voice changer</h1>\n<p>Next is the introduction of voice changer.</p>\n<h2>RVC-webUI</h2>\n<p>RVC has become a hot topic recently, but this is an operation over the webui.<br>\nddPn08's RVC-fork can convert audio using webAPI while running RVC-webUI locally. If you are familiar with RVC, please try it.<br>\n<a href=\"https://github.com/ddPn08/rvc-webui/tree/main\" target=\"_blank\">https://github.com/ddPn08/rvc-webui/tree/main</a></p>\n<h2>VoRAS</h2>\n<p>Most of RVC converts to one speaker with one model, but in my library called voras, which was created by remodeling RVC, speech conversion to more than 2000 people trained using LibriTTS-R can be performed at once. It is possible by simply switching the input id in one model. Use it for data augmentation.<br>\n<a href=\"https://github.com/nadare881/voras-webui-beta/tree/main\" target=\"_blank\">https://github.com/nadare881/voras-webui-beta/tree/main</a></p>\n<p>voras_pretrain_libritts_r.pth is a pretrain model trained with over 2000 speakers.<br>\n<a href=\"https://huggingface.co/datasets/nadare/voras/tree/main\" target=\"_blank\">https://huggingface.co/datasets/nadare/voras/tree/main</a></p>\n<p>This is a demonstration of speech conversion using a model trained by five Japanese free sound sources, but it is also possible to convert speech through webAPI.<br>\n<a href=\"https://twitter.com/i/status/1682371135718199296\" target=\"_blank\">https://twitter.com/i/status/1682371135718199296</a></p>\n<p>Voras plans to leave it without any fixes other than bugfixes at least until the end of the competition. If you have any questions about how to use voras ask me in this thread. (I don't know if I see them often)</p>",
      "rawMarkdown": "Hello Kaggler\n\nI will introduce free software that seems to be useful for speech recognition.\n\n# AI voice remover\nThe voice of the training data for this competition contains some voices other than human voices.\nFor this, you can separate the environmental sound and the human voice by using the ultimate vocal remover.\nhttps://ultimatevocalremover.com/\n\n# AI voice changer\nNext is the introduction of voice changer.\n\n## RVC-webUI\nRVC has become a hot topic recently, but this is an operation over the webui.\nddPn08's RVC-fork can convert audio using webAPI while running RVC-webUI locally. If you are familiar with RVC, please try it.\nhttps://github.com/ddPn08/rvc-webui/tree/main\n\n## VoRAS\nMost of RVC converts to one speaker with one model, but in my library called voras, which was created by remodeling RVC, speech conversion to more than 2000 people trained using LibriTTS-R can be performed at once. It is possible by simply switching the input id in one model. Use it for data augmentation.\nhttps://github.com/nadare881/voras-webui-beta/tree/main\n\nvoras_pretrain_libritts_r.pth is a pretrain model trained with over 2000 speakers.\nhttps://huggingface.co/datasets/nadare/voras/tree/main\n\nThis is a demonstration of speech conversion using a model trained by five Japanese free sound sources, but it is also possible to convert speech through webAPI.\nhttps://twitter.com/i/status/1682371135718199296\n\nVoras plans to leave it without any fixes other than bugfixes at least until the end of the competition. If you have any questions about how to use voras ask me in this thread. (I don't know if I see them often)",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2360698": "Hello Kaggler\n\nI will introduce free software that seems to be useful for speech recognition.\n\n# AI voice remover\nThe voice of the training data for this competition contains some voices other than human voices.\nFor this, you can separate the environmental sound and the human voice by using the ultimate vocal remover.\nhttps://ultimatevocalremover.com/\n\n# AI voice changer\nNext is the introduction of voice changer.\n\n## RVC-webUI\nRVC has become a hot topic recently, but this is an operation over the webui.\nddPn08's RVC-fork can convert audio using webAPI while running RVC-webUI locally. If you are familiar with RVC, please try it.\nhttps://github.com/ddPn08/rvc-webui/tree/main\n\n## VoRAS\nMost of RVC converts to one speaker with one model, but in my library called voras, which was created by remodeling RVC, speech conversion to more than 2000 people trained using LibriTTS-R can be performed at once. It is possible by simply switching the input id in one model. Use it for data augmentation.\nhttps://github.com/nadare881/voras-webui-beta/tree/main\n\nvoras_pretrain_libritts_r.pth is a pretrain model trained with over 2000 speakers.\nhttps://huggingface.co/datasets/nadare/voras/tree/main\n\nThis is a demonstration of speech conversion using a model trained by five Japanese free sound sources, but it is also possible to convert speech through webAPI.\nhttps://twitter.com/i/status/1682371135718199296\n\nVoras plans to leave it without any fixes other than bugfixes at least until the end of the competition. If you have any questions about how to use voras ask me in this thread. (I don't know if I see them often)"
  }
}