Get a professionally trained, custom AI voice model built from your own recordings — with source audio cleaning included so your model sounds its absolute best from the start.
I Will Train a Custom RVC V2 or GPT-SoVITS AI Voice Model
Custom RVC V2 model trained on a short dataset with full source audio cleaning.
- RVC V2 voice model training
- Dataset up to 5 minutes of recorded audio
- Professional source voice audio cleaning
- Custom voice training data preparation
- Delivered as a ready-to-use model file
- Suitable for speech-to-speech conversion
Extended dataset training with singing vocal support — ideal for music and AI cover projects.
- RVC V2 voice model training
- Dataset up to 10 minutes of recorded audio
- Professional source voice audio cleaning
- Singing vocals training included
- 1 revision included
- Suitable for AI covers, music production and speech synthesis
Full-scope training with singing vocals, longer dataset support and maximum refinement.
- RVC V2 or GPT-SoVITS voice model training
- Dataset beyond 10 minutes of recorded audio
- Professional source voice audio cleaning
- Singing vocals training included
- 1 revision included
- Priority handling with thorough fine-tuning pass
Request a Custom Offer
Log In to Request a Custom Offer
Create a free account or log in to request a personalised offer from this Zinner.
Log In / RegisterAsk a Pre-Sale Question
Log In to Ask a Question
To reduce platform spam, pre-sale messages can only be sent by logged-in users.
Create a free account or log in to message this Zinner directly.
Log In / RegisterAt a Glance
Key details about this service to help you decide. Generated by Zinn Hub, not the seller.
Value Position
Model Type
Dataset Requirement
Singing Support
Pre-Training Cleanup
What You'll Receive
Full Description
Whether you want to clone your own voice, recreate a character, or build a unique synthetic speaker, a well-trained RVC V2 voice model is the foundation everything else depends on. This service handles the full training pipeline for you — from cleaning your raw audio through to delivering a ready-to-use model file.
Based in London, England, Zinn Digital specialises in custom AI voice model training using RVC V2 and GPT-SoVITS. Every order includes professional source audio cleaning before training begins, which removes background noise, inconsistencies and artefacts that would otherwise degrade the final model quality. You simply provide your voice recordings; the technical heavy lifting is done for you.
The entry tier covers datasets of under 5 minutes — ideal for quick voice clones, testing a concept, or straightforward speech-to-speech conversion use cases. Stepping up to the Standard tier extends coverage to longer datasets and unlocks singing vocal training, making it suitable for music productions, AI covers and vocal synthesis projects. The Full tier provides the same expanded scope with additional revision flexibility and the maximum delivery window to allow thorough fine-tuning.
RVC V2 models are used for speech-to-speech conversion: you generate speech via any text-to-speech system, then run it through the trained model to output audio in your target voice. For local use, a compatible GPU (NVIDIA GTX 1000 series or newer, or AMD RX 6000 series or newer with at least 4 GB VRAM) and around 10 GB of free storage are recommended. The model itself is language-agnostic — it can be trained on recordings in any language, though cross-language conversion may show some performance variation.
A dataset of around 10 minutes of clean, consistent audio tends to produce strong results. The cleaner and more consistent your recordings, the better the output — which is exactly why source audio cleaning is included as standard on every tier.
This service is right for music producers creating AI vocal covers, developers building voice-enabled applications, content creators wanting a consistent synthetic speaker, and anyone exploring custom voice synthesis for creative or commercial projects.
Once your order is placed, simply share your dataset via the order chat and the work begins straight away.
Zinner Quality Guarantee
Every Zinner is reviewed and approved before joining the platform.
All services are backed by our quality assurance commitment.
Your payment is protected until you approve the delivered work.
Compare Packages
| Feature | Basic | Standard | Full |
|---|---|---|---|
| Delivery Time | 1 days | 2 days | 3 days |
| Revisions | 0 | 1 | 1 |
| RVC V2 voice model training | ✓ | ✓ | ✕ |
| Dataset up to 5 minutes of recorded audio | ✓ | ✕ | ✕ |
| Professional source voice audio cleaning | ✓ | ✓ | ✓ |
| Custom voice training data preparation | ✓ | ✕ | ✕ |
| Delivered as a ready-to-use model file | ✓ | ✕ | ✕ |
| Suitable for speech-to-speech conversion | ✓ | ✕ | ✕ |
| Dataset up to 10 minutes of recorded audio | ✕ | ✓ | ✕ |
| Singing vocals training included | ✕ | ✓ | ✓ |
| 1 revision included | ✕ | ✓ | ✓ |
| Suitable for AI covers, music production and speech synthesis | ✕ | ✓ | ✕ |
| RVC V2 or GPT-SoVITS voice model training | ✕ | ✕ | ✓ |
| Dataset beyond 10 minutes of recorded audio | ✕ | ✕ | ✓ |
| Priority handling with thorough fine-tuning pass | ✕ | ✕ | ✓ |
Portfolio
Examples of the seller's work related to this Zinn.

Train a Custom RVC V2 or GPT-SoVITS AI Voice Model


Train a Custom RVC V2 or GPT-SoVITS AI Voice Model

Extra Information
Why Choose Me
Tools I Use
Perfect For
Frequently Asked Questions
Around 10 minutes of clean, consistent audio tends to produce strong results. The Basic tier covers recordings under 5 minutes. For best quality, aim for clear speech with minimal background noise and a consistent microphone distance throughout.
Not directly — RVC V2 models are designed for speech-to-speech conversion rather than text-to-speech. The typical workflow is to generate speech using any text-to-speech tool (with any voice), then run that audio through your trained RVC model to convert it into your target voice. This gives you full control over the output.
Yes, the RVC V2 model can be trained on recordings in any language — there are no language restrictions for training. If you plan to convert speech from one language into another, be aware there may be some performance degradation in cross-language conversion scenarios.
To run the model on your own machine you will need a GPU with at least 4 GB of VRAM and approximately 10 GB of free storage. For NVIDIA, a GTX 1000 series card or newer is suitable. For AMD, an RX 6000 series card or newer is recommended.
You need to supply your voice recordings as the dataset. The cleaner and more consistent the audio, the better the trained model will perform. Source audio cleaning is included in every tier, but recordings free from heavy background noise will always yield the best results. Simply share your files via the order chat once your order is placed.
Before training begins, the provided audio is processed to remove background noise, artefacts and inconsistencies that could negatively affect the model. This step is included as standard on every tier and is a key part of ensuring the final voice model is as clean and accurate as possible.
The trained model will be delivered as a ready-to-use model file via the order manager. You will receive instructions on how to load and use it, including recommended local software for running speech-to-speech conversion.
Both are AI voice model frameworks used for voice cloning and synthesis. RVC V2 is widely used for speech-to-speech conversion and AI vocal covers. GPT-SoVITS can offer different performance characteristics particularly for certain languages and use cases. If you have a preference or a specific use case in mind, mention it in the order requirements and the most suitable framework will be selected.
Customer Reviews
See what our customers say about this Zinn
Great work! The model sounds authentic, airy and true to Leslie Cheung’s tone. Thanks, but I don't know why you did it so fast (1 day). 3500 epochs, thanks.
Perfection! Well understanding what my requirements were!
Same-day service was provided. The model has no artefacting and exceeds my expectations. A recommendation to get the most out of your purchase is to make sure you use w-okada and not a third-party service like voice.ai, as they often won't take into account index files, which are vital.
Very quick!
Very professional, very prompt, and always responsive to requests. I will continue to cooperate.
Only logged in customers who have purchased this product may leave a review.









