OpenClaw AI Video Editor
Agentic AI video editor — autonomous agent that plans and executes natural-language video edits end-to-end: viral clips, auto captions, vertical 9:16 reframe, chroma key, silence removal, motion tracking, B-roll, voiceover, music generation, and MP4 / multi-platform export through a single agentic tool.
Install
openclaw plugins install clawhub:openclaw-ai-video-editor
OpenClaw AI Video Editor
The agentic production environment for video. Levea combines probabilistic multimodal intelligence with a deterministic production harness to plan, execute, verify, and revise complete video edits from natural-language instructions.
Product Distinction: Unlike one-shot video generators, Levea maintains an editable project containing the timeline, assets, layout, animation, audio, brand rules, and delivery requirements. Generative models become tools inside the workflow—not the owner of the workflow.
🧠 System Architecture: Probabilistic Intelligence + Deterministic Video-Production Harness
Levea applies the agentic system pattern to video: probabilistic reasoning operating through typed tools, persistent state, deterministic execution, and automated verification.
Creative Intent + Source Media
│
▼
Probabilistic Multimodal Intelligence (Frontier Models)
│
▼
CanonicalActionIR (179 Typed Actions)
│
▼
CanonicalPlanCompiler
│
┌─────────────┴─────────────┐
▼ ▼
Workflow DAG Scene IR
(Task Dependencies & Gating) (Timeline & Layer Hierarchy)
└─────────────┬─────────────┘
│
▼
Deterministic Video-Production Harness
├── Perception Engine (Face Tracking, Active Speaker, Shot Cuts)
├── Timeline and Scene Graph (Tracks, Layers, Geometry)
├── Media Operators (Silence Cuts, Slip/Slide, Ripple, Split-Screen)
├── Caption & Typography Engine (39 MSDF Fonts, 41 Templates, Kinetic Word Sync)
├── Animation & Motion Physics (Damped Spring Solvers, 17 Blend Modes)
├── Audio Mastering (EBU R128 Loudness, Google Cloud TTS, Auto-Duck)
├── Composition & Asset Execution
│ ├── HyperFrames — procedural charts, audio waveforms, cards, 3D
│ ├── Vulkan/WebGPU — GPU shaders, primitives, final compositing
│ ├── Lottie — authored vector animations
│ └── Veo / Imagen 3 — generated supporting media
├── Gated Executor (gatedExecute Atomic Mutations)
├── Project Versioning (Immutable Scene History)
└── Export Pipeline (Hardware-Accelerated / Vulkan MP4)
│
▼
Verification → Bounded Repair → Export
1. The Probabilistic Multimodal Intelligence Layer
Levea’s planning layer combines frontier multimodal models and reasoning models with domain-specific editing policies, structured media representations, and constrained tool execution. This layer operates probabilistically to handle high-level creative synthesis:
- Planning & Layout Intents: Translating raw prompts, transcripts, and visual assets into a structured, executable task graph.
- Linguistic & Semantic Segmenting: Running part-of-speech analysis to divide spoken lines dynamically into natural phrase breaks, setup eyebrow clauses, and punchline highlights.
- Model-Agnostic Orchestration: Serving as a flexible, model-agnostic orchestration layer across frontier models (such as Gemini, Claude, and OpenAI) and specialized media-generation models.
2. The Typed Edit Graph / Media IR
The strategic core of Levea is its intermediate representation (IR)—a structured, serializable Abstract Syntax Tree (the Scene Graph) representing your editing state. This Typed Edit Graph is the foundation of Levea's editing moat, enabling:
- Editability & Replayability: Re-rendering and modifying specific elements without needing to regenerate the entire video from scratch.
- Project Versioning & Diffs: Enabling Git-like, non-destructive timeline commits and seamless Undo/Redo cycles.
- Model Portability & Verification: Decoupling the creative model's intent from the final render, allowing automated safety verifications before compiling.
3. The Deterministic Video-Production Harness
A common misconception is to define a "harness" too narrowly as only the renderer. A true harness is the entire deterministic software environment that surrounds the AI model, giving its probabilistic reasoning persistent state, structured tools, automated verification, and repeatable execution:
- Timeline and Scene Graph: Maintains the state-of-truth and track-relative alignments.
- Media Operators: Standardized, repeatable functions to cut silences, sync multi-cam tracks, and crop visual frames.
- Caption and Layout Engine: Advanced text shaping (
cosmic-text/ HarfBuzz) and real-time bounding box constraints (avoiding faces and social media safe zones). - Animation System & Renderer (Vulkan / WebGPU): Computes smooth, 60 FPS velocity-proportional motion blur speed-streaks (
velocity.rs) and compiles hardware-accelerated layouts. - Validators & Versioning: Structural verification loops that run final safe-zone checks and track timeline versioning.
- Export Pipeline: Handles multi-platform packaging (FFmpeg, NVENC, H.264) for final delivery.
Operating Model Comparison
| Workflow Aspect | Traditional Editing | Levea Agentic Production |
|---|---|---|
| User Interaction | Human directly manipulates the timeline | Agent manipulates a typed project on the user’s behalf |
| Workflow Coordination | User coordinates individual, isolated tools | Agent plans and orchestrates the complete workflow |
| Primary Output | Final exported flat file is the sole output | Editable project state and rendered media are both outputs |
| Quality Control | Quality control and bounds checks are manual | Automated verification checks run before delivery |
| Revisions | Revisions require manually reopening the timeline | Natural-language revisions modify persistent project state |
Contents
- Quickstart — running in 60 seconds
- What it does — say it, get it
- Capability surface — the full toolset
- Goals & briefs — from one line to a full creative brief
- API reference — endpoints, request/response, streaming, async
- Who it's for · Safety · Media
- Links & channels
Quickstart
The portable interface is the levea-mcp-server MCP server — one server, every MCP client (Claude Desktop, Claude Code, Cursor, Cline, OpenClaw, Hermes), one tool surface, one backend contract, so nothing drifts per platform.
1. Get an API key — sign up at livecore.ai and generate an OpenClaw API key.
2. Add the MCP server — the same npx line works for every MCP client:
{
"mcpServers": {
"levea": {
"command": "npx",
"args": ["-y", "levea-mcp-server"],
"env": {
"LEVEA_API_URL": "https://api.livecore.ai",
"LEVEA_API_KEY": "your-key-from-livecore.ai"
}
}
}
}
3. Use it — describe the desired outcome. Levea plans, edits, verifies, and exports the video—without requiring manual timeline work, keyframing, or plugin orchestration.
Per-client setup
| Client | How |
|---|---|
| Claude Desktop | Add to ~/Library/Application Support/Claude/claude_desktop_config.json |
| Claude Code | Run claude mcp add levea npx -y levea-mcp-server |
| Cursor | Add in Settings → Features → MCP as a standard command tool |
| Cline | Configure in settings or add to your Cline MCP servers list |
| OpenClaw | Install via terminal: clawhub install ai-agentic-video-editor |
What it does
Levea produces editable, verifiable video projects in minutes rather than returning an opaque generated clip. Here are some examples of what you can accomplish:
- TikTok / Reels Prep: "Make this clip vertical, remove silences, and add bold captions."
- Viral Clip Batching: "Turn this long video into five 15-30 second viral clips."
- AI Green-Screen Replacement: "Key out the green screen, add a studio background, and keep the speaker centered."
- Contextual B-Roll Placement: "Add B-roll over the product mention and duck the music under speech."
- Multi-Cam Angle Switching: "Sync the podcast angles and cut between cameras based on the active speaker."
- Brand Styling: "Add motion graphics, lower thirds, stat callouts, and brand colors."
- Compliance & Redaction: "Blur background faces, hide the license plate, repair safe zones, and run a final delivery check."
Capability surface
The full toolset behind the prompt. Generative models act as engines inside Levea—not the owners of the workflow:
🟢 Production-Ready
- Timeline State & Layers — insert, update, replace, or delete layers (video, audio, image, text, shape, solid, adjustment, group, light, vfx, lottie); track-relative alignments and nesting.
- Caption & Text Engines — auto-generate captions from the transcript; 41 built-in templates (Hormozi, Minimal-Pro, Karaoke, Typewriter...); word timing and keyword emphasis; shape text along curved paths; dual-font visual contrast.
- Audio & Music — parametric EQ, loudness (EBU R128 LUFS) normalization, EQ presets, crossfades, background music loops, auto-ducking, stem separation (Demucs), AI music generation, and Google Cloud TTS voiceover with optional voice cloning (Chirp 3 Instant Custom Voice).
- Color Grading & Compositing — LUT filters, lift/gamma/gain wheels, neural matting (Robust Video Matting), chroma key, mask blocks, vignettes, and LaMa content-aware inpainting for on-screen text / logo / watermark removal.
- Privacy & Compliance — face blur, privacy redaction, and profanity muting.
- Perception & Motion — multi-cam sync, active-speaker detection and cuts, face-aware motion tracking with automatic zoom-follow framing, scene/shot cut analysis, cross-clip face identity search, and on-screen text region detection with OCR recognition (PP-OCRv5 det + rec).
- AI B-Roll & Generative Media — prompt-driven B-roll generation, generative voiceovers, and AI image/video asset generation wired as canonical actions.
- High-Level Presets — vertical reframing, viral clips generation, active-word keyword emphasis, and silence cleanup.
🟡 Model- or Deployment-Dependent
- Quality and availability of generated video, images, B-roll, music, voiceover, and voice cloning vary by model tier and account quota (all are wired as canonical actions in production).
- Neural alpha matting and background replacement quality.
- CJK / RTL caption rendering — cosmic-text bidi shaping and RTL language handling ship in production; the 39-font MSDF atlas is Latin-ASCII, so CJK/RTL glyph coverage depends on system-font fallback when
preserveUnicodeis enabled.
Goals & briefs
autonomous_edit accepts anything from a five-word command to a five-hundred-word creative brief.
1. Direct commands — a single edit, single result. The What it does examples above are all direct commands.
2. Open-ended goals — hand the agent a goal instead of a command; it inspects the asset, then proposes a plan:
"Watch this and propose edits to make it more engaging for TikTok"
"Look at the first 30 seconds and suggest 3 ways to hook the viewer"
"Review this footage and propose edits to tighten the pacing"
Pair these with requirePlanApproval: true for propose-only mode — the agent stops after planning and waits for your approval before execution.
API reference
Base URL: {LEVEA_API_URL} (production: https://api.livecore.ai). Auth on every request:
Authorization: Bearer {LEVEA_API_KEY}
All paths are under /api/v1/misc/openclaw. For complete internal schemas, payload contracts, deduplication rules, and response objects, please refer to AGENTS.md.
Core endpoints
| Verb | Path | What it does |
|---|---|---|
POST | /v1/execute | Primary endpoint. Executes a prompt and returns the results. |
POST | /v1/execute_streaming | SSE stream of step-by-step progress updates. |
POST | /v1/queue-edit | Enqueues background render or generation jobs. |
GET | /v1/jobs/{jobId} | Polls async job or render statuses. |
Progressive Execution Visibility (SSE Streaming)
The /v1/execute_streaming endpoint returns a Server-Sent Events (SSE) stream of structured events representing the editor's live execution progress. These events expose useful API concepts:
plan_step: Details the structured task list compiled by the DAG layout planner.progress: Numeric execution progress percentage.decision_summary: Concise structural rationale behind tool selections.execution_update: Specific real-time updates of running media operators.
Links & channels
| Asset | Link |
|---|---|
| Studio & API keys | livecore.ai |
| API base | https://api.livecore.ai |
| npm package | levea-mcp-server |
| GitHub Repository | agentic-ai-video-production |
Support — setup help, integration questions, or issue reports: brajendrak00068@gmail.com
