If you have ever wondered exactly how to make a VTuber model from a blank screen to a streaming-ready avatar, this guide covers the entire pipeline in one place. By the end you will know how to choose between a 2D and 3D model, how to rig it, how to add live face and hand tracking, how to wire it into OBS, and how to give your character a consistent voice that matches the avatar. Both halves matter: viewers remember a VTuber by how the model looks and how it sounds.
This is a practical, numbered walkthrough aimed at beginners who want to ship a real model rather than read theory. If you are still deciding whether VTubing is for you at all, the how to become a VTuber guide is a good starting point, and the how to make a VTuber avatar article goes deeper on the DIY tooling.
TL;DR
- Pick a path first: 2D (Live2D) for expressive anime art, or 3D (VRoid Studio) for the fastest free start.
- DIY vs commission vs maker tools: VRoid Studio is the easiest free maker; commissioning an artist costs more but saves time.
- Rigging turns art into a puppet - VRoid auto-rigs 3D, while Live2D needs manual deformer and physics work.
- Tracking software (VTube Studio for 2D, VSeeFace for 3D) maps your webcam or phone camera to the avatar.
- OBS captures the avatar with a transparent background so you can layer overlays and go live.
- Voice matters too: a real-time voice changer like VoxBooster keeps your character voice consistent and feeds OBS through a virtual mic.
What Is a VTuber Model?
A VTuber model is a digital avatar - 2D or 3D - that mirrors a streamer’s face and body movements in real time so the on-screen character speaks and emotes as the person behind it does. The model itself is just artwork plus a rig; separate tracking software reads your camera and animates it live. VTuber comes from “virtual YouTuber,” a format that started in Japan and now spans every major streaming platform. You can read background on the format on Wikipedia’s VTuber page.
In practice, “the model” is a bundle: the character art, a skeleton or set of deformers (the rig), a list of facial expressions (blend shapes), and physics for hair and clothing. Your tracking app loads that bundle and drives it with your face. Your job in building a model is to produce that bundle, then connect it to tracking and OBS.
2D vs 3D VTuber Models: Which Path Should You Choose?
The first real decision is dimension. A 2D model is a flat illustration cut into movable layers and rigged in Live2D Cubism - it looks exactly like a hand-drawn anime character but only moves convincingly within a limited range of head angles. A 3D model is a full mesh you can rotate freely, usually built in VRoid Studio and exported as a VRM file that auto-rigs for you.
Neither is objectively better. 2D wins on artistic charm and is the look most people picture when they think “anime VTuber.” 3D wins on setup speed, free tooling, and freedom of movement (full body tracking, dancing, turning around). Beginners on a budget almost always start 3D with VRoid, then commission a 2D Live2D model later once they know they will keep streaming. Here is the side-by-side.
| 2D model (Live2D) | 3D model (VRoid / VRM) | |
|---|---|---|
| Core tool | Live2D Cubism | VRoid Studio (free) |
| Art style | Hand-drawn anime 2D | Stylized anime 3D |
| Skill needed | Drawing + rigging | Slider-based, none required |
| Time to first model | Several weeks | 3-8 hours |
| Cost (DIY) | Free tier / paid Pro | Free |
| Cost (commission) | Higher (custom art + rig) | Lower to mid |
| Movement range | Limited head angles | Full 360 rotation, body |
| Tracking app | VTube Studio | VSeeFace, VTube Studio |
| Best for | Expressive anime look | Fast free start, body tracking |
Pick the row that matches your situation and commit. The biggest mistake new VTubers make is spending weeks comparing tools instead of finishing one model.
How to Make a VTuber Model: Step by Step
The steps below cover the most common beginner route - a free 3D VRoid model - and note where the 2D Live2D workflow differs. The same high-level sequence (build, rig, track, capture, voice) applies to every path.
Step 1 - Decide How You Will Get Your Model
You have three honest options:
- DIY with maker tools. VRoid Studio (3D) or Live2D Cubism (2D). Free or low cost, full control, but you do the work. Best for learning and for tight budgets.
- Commission an artist. You pay an illustrator and a rigger to build a custom model. Highest quality and zero rigging headaches, but it costs money and takes a queue of weeks. Best once you are committed.
- Buy a pre-made model. Marketplaces sell ready-rigged models you can use immediately. Faster than commissioning, cheaper than custom, but other people may use the same base.
For your first model, the DIY maker-tool route is the fastest way to learn the entire pipeline end to end without spending anything.
Step 2 - Build the 3D Model in VRoid Studio (or Draw for Live2D)
Download VRoid Studio from its official site, install it (around 600 MB, no account required), and choose Create New Model. From there:
- Set body proportions - height, head-to-body ratio, build. A head ratio near 1:5 or 1:6 gives the classic anime look.
- Shape the face with the slider panels: eyes, nose, mouth, ears, skin tone. Spend most of your time here - the face is what viewers watch.
- Style the hair using guide-curve hair groups. Use 3-5 groups for a first model to keep the triangle count reasonable.
- Dress the avatar from the built-in outfit library, recoloring or importing custom texture PNGs as needed.
If you chose the 2D Live2D path instead, this step is different: you draw the character flat in Photoshop or Clip Studio Paint, splitting every part that needs to move onto its own layer (each eye, each eyebrow, mouth open and closed, bangs, body). A basic 2D model needs 30-60 separate layers saved as a PSD before rigging.
Step 3 - Rig the Model
Rigging is what turns static art into a controllable puppet, and it is where 2D and 3D diverge most.
- 3D (VRoid): rigging is automatic. When you export to VRM, VRoid generates the bone skeleton and standard expressions for you. You do not weight-paint anything by hand.
- 2D (Live2D Cubism): rigging is manual and is the bulk of the work. You import your PSD, place warp and rotation deformers over each part, then set keyforms for each parameter (Head X, Head Y, Eye Open, mouth shapes) at the extremes so Cubism can interpolate the in-between motion. You also add physics groups so hair and accessories swing naturally.
For a deeper look at rigging tradeoffs across tools, the how to make a VTuber avatar guide breaks down the full DIY comparison including the Blender plus UniVRM pipeline for advanced 3D artists.
Step 4 - Add Expressions and Physics
Expressions are blend shapes - preset facial states your tracking software triggers. VRM models support a standard set (Joy, Angry, Sorrow, Fun, blink, and the A/I/U/E/O mouth shapes), and VRoid generates these automatically. To add extras like a wink or a tongue-out, define custom blend shapes in VRoid’s export settings, then assign them to hotkeys in your tracking app later.
Physics give hair, ribbons, and loose clothing secondary motion so they swing when your head moves. In VRoid you configure spring groups in the Physics/Collider tab: high stiffness (around 0.8) for short hair, low stiffness (0.1-0.3) for long flowing hair. Tuning this takes a little experimentation, but the difference between rigid and physics-enabled hair is dramatic on stream.
Step 5 - Export the Model File
- 3D: in VRoid, go to Export as VRM, fill in author name and license, and save the
.vrmfile. Typical size is 20-80 MB with a triangle count of roughly 30,000-70,000 depending on hair complexity. - 2D: in Cubism, export the
.moc3plus.model3.jsonbundle along with the texture atlas. The.model3.jsonis the file your tracking software loads.
Keep this exported file backed up - it is the single most valuable asset you have built.
Step 6 - Set Up Face and Hand Tracking Software
Tracking software is the engine that brings the model to life. Load your exported file, point it at your camera, and it maps your movements onto the avatar in real time.
- VTube Studio is the standard for 2D Live2D models and also accepts 3D VRM. It uses a webcam or a phone camera (the paid iOS app uses ARKit depth sensors for the lowest-latency, most accurate tracking).
- VSeeFace is a popular free option for 3D VRM models using a standard Windows webcam. It also exposes a live blend-shape value window that is useful for debugging which parameter is not responding.
For hand tracking, some setups add a separate tracking layer (a webcam-based hand tracker or a Leap Motion-style device) that feeds finger and arm movement into the model. Hand tracking is optional - most VTubers launch with face-only tracking and add hands later.
In your tracking app, run a quick checklist: confirm blink works (adjust sensitivity if you wear glasses), test mouth sync by saying vowels out loud, and rotate your head to the extremes to look for mesh clipping at the neck.
Step 7 - Wire the Avatar Into OBS
OBS Studio is the free broadcasting software that combines your avatar, overlays, and audio into one stream. To bring the avatar in:
- In your tracking app, set the background to a solid color or enable its transparent/spout output.
- In OBS, add your tracking app as a Game Capture or Window Capture source.
- If you used a solid color background, add a Chroma Key filter on that source to make it transparent so only the avatar shows.
- Layer your scene: avatar on top, webcam overlays, alerts, and a background image or scene art beneath.
- Add your microphone (or voice changer’s virtual device) as the audio input so your voice and avatar are captured together.
Once the avatar moves cleanly over a transparent background in OBS, the visual side of your VTuber setup is done.
The Audio Side: Giving Your VTuber Model a Voice
Here is the half most guides skip. A VTuber model is only as convincing as the voice coming out of it. Viewers form an attachment to a character, and a character that sounds like a different person every stream breaks the illusion. This is why so many VTubers run their microphone through a real-time voice changer to lock in a consistent persona voice.
A real-time voice changer sits between your microphone and your streaming software. It processes your voice on the fly - shifting pitch, applying a character voice, smoothing tone, and removing background noise - then outputs the result as a virtual microphone that OBS treats like any normal mic. That means zero conflict with your avatar pipeline: VTube Studio drives the visuals, the voice changer handles the audio, and OBS captures both.
VoxBooster is built for exactly this. It runs locally on Windows 10 and 11, processes your mic in real time with low latency, and installs no kernel driver. You can shift pitch to match a younger or older character, apply an on-device AI voice clone so your persona has its own custom voice, add effects, and run noise suppression all at once. Because the voice clone runs on your machine, your character voice stays consistent every session without depending on a cloud service.
To connect it: install VoxBooster, choose or build your character voice, and select its virtual microphone as the audio input device inside OBS (and inside Discord or any game, if you want the same voice there too). For a step-by-step on the voice-changer side, see how to use a voice changer on Discord - the OBS routing concept is identical. VoxBooster’s soundboard with hotkeys is also handy for reaction sounds during a stream, and the built-in transcription can caption your VODs after the fact.
If you want to test the full audio side before committing, you can download VoxBooster and use the trial to dial in your character voice, then check the pricing options once you know it fits your setup.
Common Mistakes When Making Your First VTuber Model
A few avoidable errors trip up almost every new VTuber:
- Over-planning the model and never shipping it. Pick one path and finish a basic version. You can always upgrade later.
- Ignoring physics. Rigid, lifeless hair is the fastest tell of a rushed model. Spend twenty minutes tuning spring groups.
- Skipping the tracking test. Always rotate your head to the extremes and test every expression hotkey before going live, not on stream.
- Forgetting the voice. A great avatar with an inconsistent or noisy voice still feels unfinished. Decide on your character voice early.
- Not backing up the model file. Your
.vrmor.moc3bundle is irreplaceable work - keep copies.
FAQ
How much does it cost to make a VTuber model? A free DIY model costs nothing using VRoid Studio or the Live2D Cubism trial. A commissioned Live2D model from an artist commonly ranges from a few hundred to a few thousand dollars depending on art quality, expressions, and rigging complexity. Most beginners start free.
Can I make a VTuber model for free? Yes. VRoid Studio exports a fully rigged 3D VRM avatar at no cost, and Live2D Cubism has a free tier for simple 2D models. Tracking apps like VTube Studio and VSeeFace have free versions, so a zero-budget setup is realistic.
Is a 2D or 3D VTuber model better for beginners? 3D is faster to start with because VRoid Studio auto-rigs your model and exports a ready-to-track file in hours. 2D Live2D models look more expressive and illustrator-like but require drawing skill and weeks of rigging work, so most beginners pick 3D first.
What software do I need to make my VTuber model move? You need a tracking app that reads your webcam or phone camera and maps your face to the avatar. VTube Studio is the standard for 2D Live2D models, and VSeeFace handles 3D VRM models. Both run on Windows and have free versions.
How long does it take to make a VTuber model? A basic VRoid 3D model takes roughly three to eight hours. A polished Live2D 2D model takes several weeks of drawing and rigging. A fully custom Blender plus Unity avatar can take one to three months for someone new to 3D modeling.
Do I need a special voice to be a VTuber? No, but a consistent character voice helps your persona feel real. Many VTubers use a real-time voice changer to keep the same on-stream voice every session. A virtual microphone routes the processed audio into OBS like any normal mic input.
How do I put my VTuber model into OBS for streaming? Add your tracking app as a Game Capture or Window Capture source in OBS, then enable a chroma key or transparent background so only the avatar shows. Layer your overlays on top. Route your microphone or voice changer as the audio input.
Conclusion
Making a VTuber model is far more approachable than it looks from the outside. Choose your dimension - 3D VRoid for the fastest free start, or 2D Live2D for that hand-drawn anime look - build the art, rig it (automatic in 3D, manual in 2D), add expressions and physics, then load it into VTube Studio or VSeeFace and capture it in OBS. That is the complete visual pipeline.
The piece that separates a memorable VTuber from a forgettable one is treating the voice as seriously as the avatar. A consistent character voice, free of background noise, makes your persona feel like a real, recurring character instead of a webcam with a mask on. Pair your finished model with a real-time voice changer and you have a full identity: a face that moves with you and a voice that stays in character every single stream.
When you are ready to lock in your character voice, download VoxBooster and run through the trial - it covers everything you need to test the voice clone, effects, and noise suppression alongside your new model before you commit.