OpenFork Desktop
The local GPU worker for the AI movie studio your agent can operate.
Available Services
MiniMax-H3 joint video and synchronized stereo audio generation; experimental 6GB offload preview tier
High-quality video (Text/Image-to-Video)
LTX-2.3 22B quantized Distilled 1.1: native audio+video generation (24GB)
Official Lightricks LTX-2 Trainer package with the LTX-2.3 22B dev checkpoint. Standard upstream training targets 80GB+ VRAM; this lane targets the upstream 32GB low-VRAM INT8 config. 24GB is not an official supported trainer target.
DreamID-Omni FP8 talking-head video with identity and voice reference inputs
ControlNet & Flux Image Gen
High-quality Z-Image (Q4_K_M GGUF)
Experimental Krea 2 Turbo GGUF Q3 text-to-image via ComfyUI-GGUF with Qwen3-VL text encoder and Qwen Image VAE
NVIDIA PiD 4x image super-resolution using the Flux/Z-Image-compatible VAE decoder path.
Anima text-to-image illustration model
FLUX.1 Kontext [dev] GGUF Q4_K_M for 8GB VRAM; supports optional two-image composition by precomposing the source and identity reference
Qwen-Image-Edit-2511 instruction editing plus Qwen-Image-2512 character LoRA text-to-image inference
Qwen-Image-2512 character LoRA text-to-image inference using the Unsloth Q4_K_M GGUF diffusion model for 24GB GPUs
Ultra-fast instruction-based editing with an optional second reference image (2 steps)
Clone-ready text-to-speech
High-quality voice cloning
Alibaba's multilingual TTS with 9 premium voices
Rednote HiLab dots.tts MeanFlow-distilled continuous TTS and zero-shot voice cloning, smoke tested on the 6GB image
XML-driven expressive speech generation with zero-shot voice cloning
AI-powered music composition
Stable Audio 3 Small-SFX sound-effects-only text-to-audio
AudioX text-to-audio and video-conditioned sound generation
Local Qwen3 4B LLM for workflow planning on low-VRAM providers
4-bit quantized music generation (~8GB VRAM)
ACE-Step 1.5 XL music generation for 16GB VRAM
Video-to-audio synthesis (MMAudio small_44k)
High-efficiency speech restoration and enhancement
High-fidelity video-to-audio generation (PRiSM)
daVinci-MagiHuman through Wan2GP with Magi Human Distill SR1080 quanto int8 weights. Use for realistic portrait talking-head/lip-sync shots only. Requires a realistic portrait start image; generates synchronized speech/audio from the text prompt.
Experimental SCAIL-2 14B WanGP candidate using the DeepBeepMeep int8 SCAIL-2 weights, SAM3 Magic Mask assets, and the native scail2_14B WanGP model type. Use for OpenFork 16GB smoke testing before production runs.
Vista4D 384p49 through Wan2GP for source-video novel-view reshooting with predefined camera trajectories.
State-of-the-art video super-resolution (24GB VRAM required)
Baidu ERNIE-Image Turbo text-to-image with CPU offload for 8GB GPUs
Ideogram 4 text-to-image using NF4 weights with CPU offload and structured JSON layer prompts; true 16GB Vast smoke passed on 2026-06-09
DomainShuttle subject-driven image-to-video using Wan2.2-DomainShuttle-A14B; registered as an 80GB service because the closest published upstream requirement is Wan2.2 A14B inference
TeleStyleV2 content/style reference transfer on Qwen-Image-Edit-2509 with the TeleStyle and DMD LoRAs; upstream target is H100 80GB
Microsoft Mage-Flow 4B RL-aligned text-to-image plus Mage-Flow-Edit semantic, appearance, restoration, structure-aware, single-reference, and multi-reference editing
ShotPlan cinematic multi-shot text-to-video using its released Wan2.2 A14B high-noise planning-token checkpoint and exact user-authored hard-cut frame positions
HOMIE subject-consistent reference-to-video on Wan2.1 14B with Qwen3-VL reasoning, grouped multi-subject and multi-view inputs, abstract/logo references, and optional OCR-map references
ABot-World 0.5B LongForcing rollout from a scene frame with a timed keyboard action sequence; upstream reports 19GB VRAM on RTX 5090.