What is Rakuda?
An AI video editor that turns your voice into a vertical explainer video with captions, cuts and diagrams. Drop in a recording and the AI transcribes and proofreads it, cuts retakes and long pauses, removes noise, adds moving diagrams, real maps and screenshots of official pages that match what you say, and finishes a 1080×1920 MP4.
Can I make videos without showing my face?
Yes. All it needs is your voice. On screen are the captions and the diagrams, maps and screenshots the AI adds to match what you say.
What kind of videos is it for?
Vertical short videos where you explain something by talking: how things work, how-tos, tips, news explainers, study and exam prep, product and service intros, personal stories.
Can I use it for TikTok, YouTube Shorts and Instagram Reels?
Yes. It exports a 1080×1920 vertical MP4 you can post as is. Loudness is set to a comfortable level for short videos (−20 LUFS), and captions sit where the apps' buttons and descriptions are unlikely to cover them.
Do I need video editing experience?
No. Drop in a recording, check the transcript, and say what kind of video you want. To change something, pick the scene and ask the AI in plain words.
How is it different from making videos with Claude Code?
The same Claude makes the visuals. With Claude Code you build transcription, cutting, rendering and export yourself and check the code as you go. You also wire up maps and public APIs yourself and re-explain the look every video. Rakuda has all of that from the start, with plugins for maps and data and themes you save for the next video: drop in a recording, pick a scene and just ask. The code the AI wrote is there if you want to look.
How is it different from auto-caption and auto-cut apps?
Beyond captions and cuts, the AI adds moving diagrams that follow what you say, maps from real road data and screenshots of official pages, and finishes a complete explainer. And unlike AI that generates footage, text, maps and numbers never come out garbled or made up.
Do I need to install an app?
No. It runs in your browser. The preview runs in the browser too, so you can play it right away without waiting for an export.
What files can I use?
Audio such as m4a, mp3, wav and webm, and video such as mp4 and mov (a video keeps its picture, cut like the voice, with the visuals and captions on top). You can also record right in the browser.
Can I use an AI voice or audio that's already edited?
Yes. Choose "finished audio" when you upload: no cuts and no processing, only the loudness set for short videos. With a script, choose "word for word" and the captions follow it exactly.
Do I need a script?
No. The AI fixes mishearings from context. If you paste a script, names, technical terms and punctuation follow it exactly.
How long does it take?
Captions and cuts take a few minutes. Visuals arrive scene by scene as they are ready, and you can play what is done while the rest is built.
Can I stop it midway?
Yes. While the AI works, the chat's send button turns into Stop. What was done so far stays, and you pay only up to where you stopped.
What if I read the same sentence twice?
The AI treats it as a retake and cuts the earlier one. Even when the transcript hears the two reads as one run, it finds them from sentence length and the pause between, splits them and listens again.
Are the diagrams and numbers accurate?
The AI builds diagrams only from what you said and adds no numbers or names you didn't mention. Illustrative examples are labelled as such. Maps use real OpenStreetMap road data, screenshots are taken of the official pages, and the AI checks it got the right page. Please give it a final check yourself before you publish.
Can I use it if I can't code?
Yes. Say what you want changed and the AI rewrites the code. The code is shown for people who want to see what changed or tweak it themselves.
Can I see what the AI changed?
Changed lines are shown as a diff, the way coding tools do. If you don't like it, Undo takes it back and Redo applies it again.
How do I bring back something the AI cut?
Turn on Show cuts to reveal the retakes and noise, then ask in the chat to keep a line, or select the range on the timeline and press Keep. You can also drag a clip's edge to restore its length.
Can I fix mistakes in the captions?
Yes. After transcription the AI fixes mishearings and spelling from context. After that, double-click a transcript line or a caption in the video to edit it.
Does it add sound effects?
Yes. In time with the diagrams — a pop as an element appears, a sweep as a line grows — mixed low enough not to get in the way of your voice. It can also pick royalty-free background music that suits the talk and duck it under your voice.
Which languages does it support?
Japanese and English. Choose the video's language when you start a project; transcription, captions and everything the AI writes on screen follow it.
Is there a logo on exported videos?
When you use only the free credits and one-off credits, a small www.rakuda.studio credit appears under the captions. It isn't there on a monthly plan.
How much does it cost?
You pay in credits for what you use; one credit is about ¥1. A finished minute of video takes about 1,500 credits (about ¥1,500), so a 60-second short is about ¥1,500. Monthly plans: Light ¥2,980/month (3,000 credits), Creator ¥9,800/month (11,000 credits), Pro ¥29,800/month (36,000 credits). You can also buy credits on their own (¥3,000 and ¥10,000). Prices include tax. See the pricing page for details.
Can I try it for free?
Yes. When your account is set up you get 1,500 free credits (about one minute of finished video) — enough to make a 60-second short and judge the result.
How are credits used?
Only when the AI works: transcription, editing, rebuilding a visual, chatting with the AI. You see an estimate before it starts and the running cost while it works. Press Stop and you pay only for what was done. Credits never expire.
Plans or one-off credits?
A plan brings credits every month, cheaper per credit on the bigger plans (up to about 21% off), and exports without the watermark. One-off credits are a single purchase for when you run short. You can change or cancel a plan at any time.
How does it compare with hiring an editor?
Outsourcing a short with captions and sound effects typically costs ¥5,000–20,000 and takes a few days. With Rakuda a 60-second short is about ¥1,500, ready in minutes, with moving diagrams and real maps included.
Can I start right away?
Demand is high, so we're inviting people from the waitlist in order. When your invite email arrives, sign in and start.