JD JoyAI
JoyAI-VL-Interaction is an open real-time video-language interaction system from JD's Joy Future Academy. The repository releases an 8B-scale vision-first interaction model, training recipe, time-aligned interaction data and a complete deployable stack. It is designed to continuously watch webcam or livestream input, decide every second whether to speak, stay silent or delegate, and respond in under a second when a visual event matters. The system is built around JoyAI-VL-8B, WebRTC streaming, vLLM/vLLM-Omni infrastructure, optional ASR and TTS services, and a background-agent service for harder subtasks. The README says the behavior is learned from more than 4M time-aligned clips and refined with reinforcement learning.
Editorial verdict
Researchers and product teams exploring proactive webcam, livestream, monitoring, cooking guidance, live commentary and event-driven visual interaction systems.
Avoid it if you need a polished hosted assistant, low-resource mobile deployment or a general offline video captioning model only.
JoyAI-VL-Interaction deserves a separate profile because it releases a full real-time interaction model and system stack, not just a static VLM checkpoint.
Apache-2.0 open-source model, dataset and deployable system; self-hosted GPU inference required
GitHub repository, Hugging Face model download, Hugging Face dataset download, Self-hosted GPU inference
Commercial use should follow the current product, API, model license and billing terms.
Review prompt, file, media upload, retention and training-use terms before sensitive workloads.
Watch a camera or livestream and speak up when a visually important event occurs.
Use the system for live game commentary, cooking guidance, stream comments or other timing-sensitive video interactions.
Use the released model, data and recipe to study proactive vision-language behavior and time-aligned interaction training.
Model names, quotas, release status, regional access and commercial terms can change quickly; recheck official sources before procurement or production use.
qwen
qwen-agentworld
gemini
doubao-ark
omnihuman-video
official · en · verified 2026-08-04
Confirms the 2026-06-20 full open-source release, 8B model, deployable stack, time-aligned data, vLLM-Omni support, sub-second latency positioning and Apache-2.0 license.
docs · zh · verified 2026-06-25
Provides the Chinese documentation mirror for positioning, quickstart, architecture, evaluation and roadmap.
benchmark · en · verified 2026-06-25
Linked from the official README as the technical report for real-time vision-language interaction intelligence.
other · en · verified 2026-06-25
Official README links this as the model release path.
other · en · verified 2026-06-25
Official README links this as the aligned interaction dataset release path.
Last checked: 2026-08-04