MOSS-VL-Realtime: An Open 11B Video Model That Can Watch and Answer in Real Time
MOSS-VL-Realtime technical introduction and quickstart: OpenMOSS has open-sourced an 11B realtime video understanding model with a 256K context window, timestamped frame input, streaming interaction, proactive silence, and dynamic correction.
Published 77 days ago. Content may be outdated.
OpenMOSS has open-sourced MOSS-VL-Realtime, the realtime streaming checkpoint in the MOSS-VL family. It is not a traditional offline video-language model that first consumes a complete video and then answers questions. It is closer to a multimodal observer: frames keep arriving, the model keeps watching, generates when needed, and waits for more evidence when the scene is still unclear.
That distinction matters. Many video-language models are built around the workflow of “upload a video, ask a question.” MOSS-VL-Realtime is designed for a different interaction pattern: the video is still happening, questions can arrive at any moment, and the answer can change as the scene changes.
30-Second Summary
- Model size: 11B vision-language model
- License: Apache-2.0, suitable for commercial use
- Context length: 256K tokens
- Core capability: continuous video-stream understanding and real-time Q&A
- Realtime behavior: proactive silence, dynamic correction, streaming output
- Input format: PIL-compatible image frames plus absolute timestamps
- Default video setup: 1 FPS, up to 256 frames, BF16 weights
- Open checkpoints: Realtime, Instruct, and Base models are available
- Best-fit use cases: video monitoring, livestream analysis, robot vision, classroom or meeting assistants, and realtime multimodal agents
How It Differs From Ordinary Video Understanding Models
Most video understanding models work in offline mode: sample frames, encode the entire video, then answer a question. That is useful for summarization, retrieval, and batch video QA, but it is not ideal when the event is still unfolding.
MOSS-VL-Realtime is built for continuous video streams.
It can start working before the video ends. You can ask a question at any timestamp, and the model answers based on the frames observed so far. If new frames overturn an earlier interpretation, it can revise the response. If the evidence is insufficient, it can choose to stay silent and keep watching.
That ability to not answer is more important than it sounds. In realtime systems, a confident guess can be worse than silence. For security, robotics, driving assistance, industrial inspection, and similar scenarios, the model should wait for more visual evidence instead of generating text just to say something.
Key Features
1. Watch While Answering
MOSS-VL-Realtime receives frames as a stream. Each frame can carry a timestamp, and the model maintains visual context inside a realtime session while producing text only when needed.
This makes it a strong fit for:
- livestream scene explanation
- video-monitoring event alerts
- robot first-person visual observation
- realtime screen or camera assistants
- the visual perception layer of multimodal agents
2. Ask Questions at Any Moment
Users do not have to wait for the video to finish. As long as the session is running, push_prompt() can inject a question into the current stream.
This interaction pattern is much closer to real usage. You are not handing a complete video to the model and waiting for a static answer. You are asking follow-up questions while the model keeps observing: “What just changed?”, “What did the person pick up?”, or “Is there any abnormal movement in the frame?”
3. Proactive Silence
The model card explicitly mentions that MOSS-VL-Realtime can emit <|silence|> when there is no meaningful visual update or when the available context is not enough.
This is a practical feature. It turns the model from something that must answer every time into something that can speak only when there is signal. For realtime monitoring, automatic narration, or agent observation streams, this can significantly reduce noise.
4. Dynamic Correction
Realtime video understanding is hard because early frames are often incomplete. A person may begin reaching for something, but only later frames reveal whether it is a cup, a phone, or something else.
MOSS-VL-Realtime is designed not to lock itself into the first interpretation. It can update its response as new frames arrive, which is closer to how humans watch a scene.
5. Timestamp-Aware Video Encoding
Each input frame can be associated with an absolute timestamp. This helps the model reason not only about frame order, but also about when an event happens, how long it lasts, and how the scene evolves over time.
That matters for tasks such as:
- event ordering
- action duration estimation
- visual-change pacing
- temporal localization in multi-turn video QA
6. XRoPE: One Coordinate Space for Text and Visual Patches
MOSS-VL uses Cross-attention Rotary Position Embedding, or XRoPE. In simple terms, it maps text tokens and visual patches into a unified three-dimensional coordinate space: time t, height h, and width w.
This gives the model a consistent positional representation across images, offline videos, and realtime streaming video. For realtime video, the model needs to know not only what appears in the scene, but also when and where it appears.
Model Configuration
| Item | Value |
|---|---|
| Model | MOSS-VL-Realtime |
| Parameters | 11B |
| Tensor type | BF16 |
| Context length | 256K |
| Vision patch size | 16 |
| Temporal patch size | 1 |
| Default video FPS | 1.0 |
| Default max video frames | 256 |
| Realtime input format | PIL-compatible image plus timestamp |
| Session scope | One active realtime session per model instance |
| License | Apache-2.0 |
Which MOSS-VL Checkpoint Should You Use?
OpenMOSS did not release only one checkpoint. The MOSS-VL family is split by usage:
| Model | Best Use |
|---|---|
| MOSS-VL-Realtime | Realtime video-stream interaction and watch-while-answering workflows |
| MOSS-VL-Instruct | Offline image/video instruction following and ordinary multimodal QA |
| MOSS-VL-Base | Continued pretraining, fine-tuning, and research |
| MOSS-VL-Instruct-0408 | Previous instruction-tuned checkpoint |
| MOSS-VL-Base-0408 | Previous base checkpoint |
If your task is image understanding, complete-video QA, or offline batch analysis, MOSS-VL-Instruct is usually the better starting point.
If your task involves cameras, livestreams, frame queues, or realtime agents, MOSS-VL-Realtime is the checkpoint to try first.
Quickstart
The minimal path is to install the MOSS-VL repository requirements and then load the Hugging Face checkpoint through Transformers.
1. Install Dependencies
git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL
conda create -n moss_vl python=3.12 pip -y
conda activate moss_vl
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
A BF16-capable NVIDIA GPU is recommended. Since the model has 11B parameters, actual memory usage depends on device mapping, FlashAttention availability, frame count, and generation length.
2. Load the Model
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
checkpoint = "OpenMOSS-Team/MOSS-VL-Realtime"
processor = AutoProcessor.from_pretrained(
checkpoint,
trust_remote_code=True,
frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
checkpoint,
trust_remote_code=True,
device_map="auto",
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
model.eval()
If FlashAttention is not available in your environment, use:
attn_implementation="eager"
3. Create a Realtime Session
The realtime capability is exposed through create_realtime_session(). You can give the model an initial instruction so it reports only meaningful visual changes instead of narrating every frame.
session = model.create_realtime_session(
processor,
initial_prompt=(
"Watch the video stream frame by frame. "
"Describe important changes only when they happen. "
"Stay silent if there is not enough evidence."
),
frame_queue_size=256,
max_tokens_per_turn=12,
max_new_tokens=4096,
do_sample=False,
)
Useful parameters:
frame_queue_size=256: bounds the pending-frame queue and helps control latencymax_tokens_per_turn=12: keeps each realtime response shortmax_new_tokens=4096: controls total generated tokens inside the sessiondo_sample=False: better for stable monitoring-style output
4. Push Video Frames
Each frame should be a PIL-compatible image and should include a timestamp.
import time
from PIL import Image
frame_paths = [
"data/frame_0001.jpg",
"data/frame_0002.jpg",
"data/frame_0003.jpg",
]
try:
session.start()
for index, path in enumerate(frame_paths):
image = Image.open(path).convert("RGB")
session.push_frame(image, timestamp=float(index))
while True:
chunk = session.poll_output(timeout=0.0)
if chunk is None:
break
print(chunk, end="", flush=True)
time.sleep(1.0)
finally:
session.close()
The timestamp=float(index) line is only a demo. In a real stream, use timestamps from the camera, player, or streaming system.
5. Ask Follow-Up Questions During the Stream
While the session is running, you can inject a question at any time:
session.push_prompt("What changed in the latest frames?")
Then drain the output:
deadline = time.monotonic() + 5.0
while time.monotonic() < deadline:
chunk = session.poll_output(timeout=0.1)
if chunk is not None:
print(chunk, end="", flush=True)
Realtime sessions stay alive waiting for future input, so production code should use a clear drain window and call session.close() explicitly.
Online vs Offline Inference
MOSS-VL-Realtime is primarily designed for online inference, meaning session-style video-stream input. It also keeps some offline image and video prompt APIs.
Think of the split this way:
- Online inference: cameras, livestreams, streaming media, robotics, realtime agents
- Offline inference: complete-video QA, batch analysis, image understanding
If you only need offline tasks, the model card suggests that MOSS-VL-Instruct is usually the preferred checkpoint.
Use Cases
Realtime Video Monitoring
The model can describe events when the scene changes and stay silent when nothing meaningful happens. That is more useful than forcing a fixed summary every few seconds.
Robot Vision Assistants
A robot’s first-person camera is a continuous stream. MOSS-VL-Realtime’s timestamped frames and dynamic correction make it a natural candidate for a robot observation module.
Livestream Understanding
Livestream analysis is not about waiting until a stream ends. It is about explaining, detecting, indexing, and reacting while the stream is still running.
Multimodal Agents
Many agents still lack realtime visual perception. MOSS-VL-Realtime can serve as a visual observer that turns video streams into textual event streams for downstream reasoning.
Industrial and Experimental Scenarios
Production lines, labs, medical devices, and screen-recording workflows all involve continuous visual streams. Proactive silence and dynamic updates can reduce meaningless output.
Limitations
MOSS-VL-Realtime is promising, but it is not a plug-and-play production system by itself.
Important constraints:
- Latency depends on hardware: GPU, frame rate, transport overhead, and decoding speed all matter
- One active realtime session per model instance: high-concurrency services need multiple instances or a scheduling layer
- Frame queues can drop old pending frames: this helps latency but can lose detail
- Control tokens require handling: tokens such as
<|silence|>,<|round_start|>, and<|round_end|>should be filtered or rendered by the application - Realtime evaluation is still evolving: the model should be judged not only by video QA accuracy, but also by response timing, silence policy, and correction quality
What Is Actually New Here?
MOSS-VL-Realtime is not interesting merely because it is another 11B video model. Its real value is that it moves video understanding from offline analysis toward realtime interaction.
The old pattern was: watch the whole video, then answer.
The new pattern is: keep watching, answer when useful, stay silent when evidence is weak, and revise when new evidence arrives.
That changes the shape of multimodal applications. Video is no longer just an uploaded file. It becomes a continuous input source. The model is no longer just a QA system; it becomes an observer.
Conclusion
MOSS-VL-Realtime pushes open realtime video understanding forward.
The most important part is not the parameter count. It is the combination of timestamp-aware frame streaming, proactive silence, and dynamic correction. Together, those features turn it from a standard video QA model into a model built for live visual scenarios.
If you are building video agents, camera assistants, livestream analysis tools, robot vision systems, or monitoring workflows, MOSS-VL-Realtime is worth trying early.
Sources:
More Articles