Chinese AI Tools
ProductsModelsIntegrationsRankingsLatest changes
Availability TrackerUse casesSubmit a toolAccount
ZHSearch

Chinese AI Tools

Independent directory for Chinese AI products. Product availability, pricing and terms can change. Verify before commercial use.

Editorial standardsClaim productUpdate infoGet featuredAdvertise

JD JoyAI

JoyAI-VL-Interaction

JoyAI-VL-Interaction is an open real-time video-language interaction system from JD's Joy Future Academy. The repository releases an 8B-scale vision-first interaction model, training recipe, time-aligned interaction data and a complete deployable stack. It is designed to continuously watch webcam or livestream input, decide every second whether to speak, stay silent or delegate, and respond in under a second when a visual event matters. The system is built around JoyAI-VL-8B, WebRTC streaming, vLLM/vLLM-Omni infrastructure, optional ASR and TTS services, and a background-agent service for harder subtasks. The README says the behavior is learned from more than 4M time-aligned clips and refined with reinforcement learning.

Globally availableFull English UIPublic APIFree

Editorial verdict

Best for

Researchers and product teams exploring proactive webcam, livestream, monitoring, cooking guidance, live commentary and event-driven visual interaction systems.

Avoid if

Avoid it if you need a polished hosted assistant, low-resource mobile deployment or a general offline video captioning model only.

Why it matters

JoyAI-VL-Interaction deserves a separate profile because it releases a full real-time interaction model and system stack, not just a static VLM checkpoint.

Trust: 5/5 sources verified, recently checkedCoverage: 100/100

Pricing

Apache-2.0 open-source model, dataset and deployable system; self-hosted GPU inference required

Payment

GitHub repository, Hugging Face model download, Hugging Face dataset download, Self-hosted GPU inference

Commercial use

Commercial use should follow the current product, API, model license and billing terms.

Privacy

Review prompt, file, media upload, retention and training-use terms before sensitive workloads.

Use-case fit

Real-time visual monitoring

Strong

Watch a camera or livestream and speak up when a visually important event occurs.

Live commentary and guidance

Strong

Use the system for live game commentary, cooking guidance, stream comments or other timing-sensitive video interactions.

Open interaction-model research

Strong

Use the released model, data and recipe to study proactive vision-language behavior and time-aligned interaction training.

Global user checklist

RegistrationConfirmedThe GitHub repository, arXiv report, blog, Hugging Face model and dataset are public.
English UIConfirmedThe repository provides English and Chinese README files and documentation mirrors.
API and docsConfirmedThe repository documents install scripts, model downloads, minimal service startup and vLLM-Omni deployment support.
International paymentConfirmedThe stack is open source; users provide their own GPU infrastructure and any external background-agent APIs.
Commercial usePartialThe repository is Apache-2.0, but production use should verify Hugging Face model and dataset cards plus privacy obligations for camera streams.
Data and privacy termsReviewWebcam and livestream inputs can contain sensitive people, homes and biometric-adjacent data; review capture, retention and consent policies before deployment.

Model names, quotas, release status, regional access and commercial terms can change quickly; recheck official sources before procurement or production use.

Pros

  • - Open 8B model, training recipe, aligned interaction data and deployable streaming stack
  • - Designed for vision-triggered proactive responses rather than turn-based prompting only
  • - Includes WebUI, real-time video inference, ASR, TTS and background-agent service layout
  • - Public docs include English and Chinese README files plus getting-started and troubleshooting guides

Cons

  • - It is a research/open-source deployment stack, not a hosted consumer assistant
  • - Real-time video interaction needs local GPU, CUDA, model downloads and streaming setup
  • - The authors note the compact 8B model does not aim to match larger commercial assistants on all open-ended chat tasks

Decision paths

qwen

qwen-agentworld

gemini

doubao-ark

omnihuman-video

Sources

JoyAI-VL-Interaction GitHub repository

official · en · verified 2026-08-04

Confirms the 2026-06-20 full open-source release, 8B model, deployable stack, time-aligned data, vLLM-Omni support, sub-second latency positioning and Apache-2.0 license.

JoyAI-VL-Interaction Chinese README

docs · zh · verified 2026-06-25

Provides the Chinese documentation mirror for positioning, quickstart, architecture, evaluation and roadmap.

JoyAI-VL-Interaction technical report

benchmark · en · verified 2026-06-25

Linked from the official README as the technical report for real-time vision-language interaction intelligence.

JoyAI-VL-Interaction Hugging Face model

other · en · verified 2026-06-25

Official README links this as the model release path.

JoyAI-VL-Interaction dataset

other · en · verified 2026-06-25

Official README links this as the aligned interaction dataset release path.

Last checked: 2026-08-04

Reviews

Availability snapshot

Availability
available
English UI
full
API
available
Rating
4.2 (0)

Latest updates

Latest changes
Open source · 2026-06-20

JD JoyAI-VL-Interaction open-source stack released

JD's Joy Future Academy has released JoyAI-VL-Interaction as an open real-time video-language interaction system. The repository includes an 8B vision-first interaction model, model weights, training recipe, more than 4M time-aligned interaction samples, a deployable streaming stack, English and Chinese docs, Hugging Face model and dataset links, and day-0 vLLM-Omni deployment support. The model is designed to continuously watch webcam or livestream input and decide when to speak, stay silent or delegate to a background model/API/agent.

Submit a review