Rakuda | AI video editor

Talk, and it becomes
an explainer video.

Drop in your voice or a filmed video. The AI cuts your retakes, adds maps, 3D, moving diagrams, real pages and papers, even background music as you talk, and finishes a vertical video with captions.

Join the waitlistDemand is high, so we're inviting people in order.

This video, too, was made from about three minutes of phone audio (in Japanese).

Done

The AI adds diagrams
that follow what you say.

The AI reads what you say, builds maps, moving diagrams, source pages and summaries, and drops each in as you say it. Below, the six scenes of the video above go from captions only to finished visuals. Click a scene to watch from there. The video is in Japanese; the AI's notes are translated.

TranscriptAs spoken (Japanese)
  1. 00:01:03
  2. 00:18:09そこで使われる基本のアルゴリズムが「ダイクストラ法」です。
  3. 00:21:18まだ決まっていない交差点の中から、一番早く着ける場所を選んで確定させる。次に、そこから伸びる道を見て、隣の交差点に着く時間を計算する。
  4. 00:47:18そこで使われるのが「A*(エースター)」です。「これまでかかった時間」に「残りの時間の目安」をプラスして、数字が小さい順に調べていきます。
  5. 01:04:21Googleは過去のデータと今の混雑状況を組み合わせて、今後の時間を予測しています。
  6. 01:28:15「地図をグラフにして、AIで重さを予測し、近い順に調べる」。
ViewerCaptions only
まず、地図はグラフという仕組みに置き換えられます。
Captions only
00:00:00
Reading scene 1 and planning its visual…
Ask the AI about this scene…↑
TimelineAI visualsCaptionsAudio

Not AI that
“generates” footage.

AI that generates whole videos tends to garble text and invent maps and numbers that only look right. Rakuda's AI builds components out of HTML.

Maps are drawn from real road data, diagrams from the steps you describe, quotes from the real page.

Below is the Summary component from the end of the video above, in English. Try editing it.

Component the AI wrotesteps · JSON
Edit it and the screen updates instantly.
ViewerSame renderer as export

Nothing to install.
Fast, in your browser.

Because the components are HTML, the preview runs right in the browser. Play without waiting for an export, and see every fix on screen at once.

  1. Open it and go

    Nothing to install. Open the browser, drop in a recording and start editing.

  2. Play without waiting

    Visuals are drawn in the browser, so it plays at once with no rendering. Jump to any scene and it's there.

  3. Changes show instantly

    A single caption letter or a component the AI rewrote — every change shows on screen right away.

Use plugins to put maps
and 3D into your talk.

The AI picks the plugins that suit what you say and builds the maps, 3D and charts. In the Plugins tab, set each to auto, on or off. The examples are real, from Rakuda's own intro video (in Japanese).

Also: background music, logos & icons, real photos

Use the maps plugin to draw the route you mention on real road data

Your voice,
or your footage.

Start the way that fits how you talk and what you have.

  1. Filmed video, kept

    A talking-head video keeps its picture, cut exactly like the voice. Visuals sit on a panel above your face; captions turn white to read over footage.

  2. Word for word

    Paste your script and the captions follow it exactly, even for audio read from it.

  3. AI narration too

    An AI voice or audio already edited goes in as finished audio: no cuts, only the loudness set.

  4. Channel themes

    Save colours, fonts and caption style as a theme, and start the next video in the same look.

What you'd build with Claude Code,
without writing code.

Rakuda's AI editor makes its visuals with the same Claude as Claude Code. The difference is everything around them: audio cleanup, retake cuts, captions, a timeline and export, plus plugins for maps and public APIs and themes you save and reuse, all there from the start.

RakudaClaude CodeEditing apps (CapCut, Vrew…)Video-generating AI
Getting startedOpen the browser, drop in a recordingBuild transcription, cutting, rendering and export yourself in a terminalInstall and goWrite a prompt
Retakes & pausesHeard and cut for you, checked against your scriptBuild the cutting yourselfSilence cut automatically; retakes found by handDoesn't edit your voice
Diagrams, maps, sourcesMade by the AI as you talk, from real maps, pages and papersPossible, scene by scene, checking the codeTemplates and stickers placed by handLook right, but text and numbers often garble
Visual styleSaved as a theme, so the next video looks the same; colors, fonts and captions are yours to changeTends to look similar every time; you re-explain it each videoPick from templatesLooks different every time
Data & mediaMaps, public APIs, papers, photos and music built in as pluginsWire up APIs and media yourselfThe app's own media libraryCan't use outside data
Reviewing & fixingPreview and timeline at once; pick a scene and askSet up your own previewBy hand on the timelineRegenerate from scratch
Coding neededNo (the code is there if you want it)YesNoNo
Export1080×1920 MP4, identical every timeDepends on your rendererMP4Different every time

It reports as it goes,
and fixes what you ask.

The AI editor proofreads the transcript, splits it into scenes and builds the visuals, reporting what it did at each step along with the cost so far and an estimate. To change something, pick the scene and just ask. The AI rewrites the component's code (settings, JavaScript, CSS) and shows the lines it changed. You can stop it at any time. Below is a real exchange from making the video above, in Japanese as it happened. Press Undo to compare with the version before.

Long pauses and retakes,
cut for you.

Stumbled? Just say it again. The AI finds the retake and cuts the earlier try, and tightens long pauses and muttering. For fine-tuning, trim, split and delete clips on the timeline. Below is a real 32-second recording (in Japanese): the red parts are cut, and “Hear it cut” plays the finished audio.

Noise gone,
a voice that's easy to follow.

It tames air-conditioner hum and room echo, lightens a boxy tone, and makes every word clear. Loudness is set to a comfortable level for short videos (−20 LUFS).

How it works

  1. Record your voice

    A phone's voice memo app or recording in the browser both work. If you stumble, just say it again. Paste a script and names and technical terms follow it too.

  2. Check the transcript

    The words that become your captions are listed. Double-click anything that's wrong to fix it. Retakes found from your script are struck through in red, and the edit cuts them.

  3. Say what kind of video you want

    Write something like "a fast-paced tech explainer", or pick a type: explainer, tips, news, story. The AI chooses the look, the title, how tight the pauses are and which plugins to use, and shows why.

  4. The AI editor edits and reports back

    It cuts retakes and pauses, splits the talk into scenes and builds the visuals, reporting what it did. The cost so far and the estimate are shown, and Stop halts it at any time.

  5. Ask for changes, then export

    Pick a scene and tell the AI what to change, and it rebuilds the visual. On the timeline you can trim, split and delete clips. Finally, export a 1080×1920 MP4.

FAQ

What is Rakuda?

An AI video editor that turns your voice into a vertical explainer video with captions, cuts and diagrams. Drop in a recording and the AI transcribes and proofreads it, cuts retakes and long pauses, removes noise, adds moving diagrams, real maps and screenshots of official pages that match what you say, and finishes a 1080×1920 MP4.

Can I make videos without showing my face?

Yes. All it needs is your voice. On screen are the captions and the diagrams, maps and screenshots the AI adds to match what you say.

What kind of videos is it for?

Vertical short videos where you explain something by talking: how things work, how-tos, tips, news explainers, study and exam prep, product and service intros, personal stories.

Can I use it for TikTok, YouTube Shorts and Instagram Reels?

Yes. It exports a 1080×1920 vertical MP4 you can post as is. Loudness is set to a comfortable level for short videos (−20 LUFS), and captions sit where the apps' buttons and descriptions are unlikely to cover them.

Do I need video editing experience?

No. Drop in a recording, check the transcript, and say what kind of video you want. To change something, pick the scene and ask the AI in plain words.

How is it different from making videos with Claude Code?

The same Claude makes the visuals. With Claude Code you build transcription, cutting, rendering and export yourself and check the code as you go. You also wire up maps and public APIs yourself and re-explain the look every video. Rakuda has all of that from the start, with plugins for maps and data and themes you save for the next video: drop in a recording, pick a scene and just ask. The code the AI wrote is there if you want to look.

How is it different from auto-caption and auto-cut apps?

Beyond captions and cuts, the AI adds moving diagrams that follow what you say, maps from real road data and screenshots of official pages, and finishes a complete explainer. And unlike AI that generates footage, text, maps and numbers never come out garbled or made up.

Do I need to install an app?

No. It runs in your browser. The preview runs in the browser too, so you can play it right away without waiting for an export.

What files can I use?

Audio such as m4a, mp3, wav and webm, and video such as mp4 and mov (a video keeps its picture, cut like the voice, with the visuals and captions on top). You can also record right in the browser.

Can I use an AI voice or audio that's already edited?

Yes. Choose "finished audio" when you upload: no cuts and no processing, only the loudness set for short videos. With a script, choose "word for word" and the captions follow it exactly.

Do I need a script?

No. The AI fixes mishearings from context. If you paste a script, names, technical terms and punctuation follow it exactly.

How long does it take?

Captions and cuts take a few minutes. Visuals arrive scene by scene as they are ready, and you can play what is done while the rest is built.

Can I stop it midway?

Yes. While the AI works, the chat's send button turns into Stop. What was done so far stays, and you pay only up to where you stopped.

What if I read the same sentence twice?

The AI treats it as a retake and cuts the earlier one. Even when the transcript hears the two reads as one run, it finds them from sentence length and the pause between, splits them and listens again.

Are the diagrams and numbers accurate?

The AI builds diagrams only from what you said and adds no numbers or names you didn't mention. Illustrative examples are labelled as such. Maps use real OpenStreetMap road data, screenshots are taken of the official pages, and the AI checks it got the right page. Please give it a final check yourself before you publish.

Can I use it if I can't code?

Yes. Say what you want changed and the AI rewrites the code. The code is shown for people who want to see what changed or tweak it themselves.

Can I see what the AI changed?

Changed lines are shown as a diff, the way coding tools do. If you don't like it, Undo takes it back and Redo applies it again.

How do I bring back something the AI cut?

Turn on Show cuts to reveal the retakes and noise, then ask in the chat to keep a line, or select the range on the timeline and press Keep. You can also drag a clip's edge to restore its length.

Can I fix mistakes in the captions?

Yes. After transcription the AI fixes mishearings and spelling from context. After that, double-click a transcript line or a caption in the video to edit it.

Does it add sound effects?

Yes. In time with the diagrams — a pop as an element appears, a sweep as a line grows — mixed low enough not to get in the way of your voice. It can also pick royalty-free background music that suits the talk and duck it under your voice.

Which languages does it support?

Japanese and English. Choose the video's language when you start a project; transcription, captions and everything the AI writes on screen follow it.

Is there a logo on exported videos?

When you use only the free credits and one-off credits, a small www.rakuda.studio credit appears under the captions. It isn't there on a monthly plan.

How much does it cost?

You pay in credits for what you use; one credit is about ¥1. A finished minute of video takes about 1,500 credits (about ¥1,500), so a 60-second short is about ¥1,500. Monthly plans: Light ¥2,980/month (3,000 credits), Creator ¥9,800/month (11,000 credits), Pro ¥29,800/month (36,000 credits). You can also buy credits on their own (¥3,000 and ¥10,000). Prices include tax. See the pricing page for details.

Can I try it for free?

Yes. When your account is set up you get 1,500 free credits (about one minute of finished video) — enough to make a 60-second short and judge the result.

How are credits used?

Only when the AI works: transcription, editing, rebuilding a visual, chatting with the AI. You see an estimate before it starts and the running cost while it works. Press Stop and you pay only for what was done. Credits never expire.

Plans or one-off credits?

A plan brings credits every month, cheaper per credit on the bigger plans (up to about 21% off), and exports without the watermark. One-off credits are a single purchase for when you run short. You can change or cancel a plan at any time.

How does it compare with hiring an editor?

Outsourcing a short with captions and sound effects typically costs ¥5,000–20,000 and takes a few days. With Rakuda a 60-second short is about ¥1,500, ready in minutes, with moving diagrams and real maps included.

Can I start right away?

Demand is high, so we're inviting people from the waitlist in order. When your invite email arrives, sign in and start.

Your next video,
just by talking.

Demand is high, so we're inviting people from the waitlist in order.
Join early.

Join the waitlist