2
3
01The task
What the next scene must respect is spread across three inputs. It only shows up when you read them together.
“Deliver medicine to grandmother.”
The bridge is broken.
A boat is waiting.
“Use the boat, and keep the medicine dry.”

02Method
Raw inputs become a grounded dependency model, which then steers every step of generation.
Turn the source, video and story into aligned states and their cross-dependencies.
What the inputs establishOnly dependencies the proposed continuation makes relevant become constraints.
c = (scope, applies-if, requires, evidence)
Plan, verify, revise, render. Then check the rendered video against the same constraints.
Plan · verify · revise · render
03Results
Overall composite score, VIS-Ground vs. the strongest baseline on each video generator.
Click a legend entry to hide a method.
Composite per backbone; the bar spans its range.
Per backbone and grounding setting, in points.
Gains concentrate in story continuity.
Composite = equal-weight mean of narrative and knowledge JCS, with 95% CI. Best per backbone in bold.
VIS-Bench
Narrative and knowledge grounding, built from CoQA passages and rendered in two styles.
SF source faithfulnessSC story continuityIF interaction fulfillmentJCS joint (mean)VC visual consistencyVQ video quality

04Examples
Each video plays the prefix, the request, then every method’s continuation, scored against five constraints.
05Try it
Real sessions from our formative study. Every question was typed by a viewer; every answer was generated live.
Pre-to-post learning gain, seven participants who passed the attention test.
Grey: one participant. Blue: the mean.
Four respondents, 1–5 scale.
06Limitations
Most outputs keep the main context. Perfect scores are rarer.
Share of outputs per score band.
Shots are synthesized separately, so voices and ambience can change between them.
The renderer draws its own text, equations and plots; our verifier checks events, not symbols.
@article{li2026visground,
title = {{VIS-Ground}: Video Interactive Storytelling with Contextual Grounding},
author = {Li, Bingxuan and Song, Yiwen and Wu, Xueqing and Pan, Yanzhou and Li, Yang and
Su, Kuang and Liu, Jingyun and Ko, Sebastian and Zhang, Huan and Zhang, Tong and
Peng, Nanyun and Pfister, Tomas and Song, Yale},
journal = {arXiv preprint},
year = {2026}
}