VIS-Ground: Video Interactive Storytelling
with Contextual Grounding

Bingxuan Li1,2 Yiwen Song1 Xueqing Wu3 Yanzhou Pan1 Yang Li1 Kuang Su1 Jingyun Liu1 Sebastian Ko1 Huan Zhang2 Tong Zhang2 Nanyun Peng1 Tomas Pfister1 Yale Song1

1 2 3
Paper Code soon
+10.3points over the strongest baseline
3/3video backbones where it ranks first
1,004viewer interventions in VIS-Bench

01The task

A request is never
just a request.

What the next scene must respect is spread across three inputs. It only shows up when you read them together.

Grounding source 𝒢

“Deliver medicine to grandmother.”

Video so far

The bridge is broken.
A boat is waiting.

Viewer asks

“Use the boat, and keep the medicine dry.”

The continuation must
  • reach grandmother
  • cross by boat
  • keep the medicine dry
Figure 1: a cat must deliver medicine to grandmother; the bridge in the video is broken, a boat is available, and the viewer asks to use the boat and keep the medicine dry. The inputs jointly induce constraints, and VIS-Ground structures the context, induces candidate-specific constraints, and generates and verifies the continuation.
Figure 1 · click to enlarge

02Method

Compile the context.
Then generate.

Raw inputs become a grounded dependency model, which then steers every step of generation.

  1. 1

    Structured Context Abstraction

    Turn the source, video and story into aligned states and their cross-dependencies.

    What the inputs establish
  2. 2

    Generation Constraints Induction

    Only dependencies the proposed continuation makes relevant become constraints.

    c = (scope, applies-if, requires, evidence)
  3. 3

    Constrained Video Generation

    Plan, verify, revise, render. Then check the rendered video against the same constraints.

    Plan · verify · revise · render
Figure 3: Overview of VIS-Ground.
Figure 3 · click to enlarge

03Results

Best on every backbone.

Overall composite score, VIS-Ground vs. the strongest baseline on each video generator.

Omni-1.1-Flash 84.2 +12.3 vs VideoGen-of-Thought
Ours
71.9
Veo 3.1 84.6 +15.3 vs Direct Generation
Ours
69.3
MiniMax-H3 78.8 +3.4 vs MovieAgent
Ours
75.4

Scores on every metric

Click a legend entry to hide a method.

Show data

VIS-Bench

The benchmark

Narrative and knowledge grounding, built from CoQA passages and rendered in two styles.

250sources 1,004interventions 2,008instances 6metrics

SF source faithfulnessSC story continuityIF interaction fulfillmentJCS joint (mean)VC visual consistencyVQ video quality

Figure 4: dataset construction and statistics.
Figure 4 · click to enlarge

04Examples

Same inputs.
Different grounding.

Each video plays the prefix, the request, then every method’s continuation, scored against five constraints.

05Try it

Pause the story.
Ask anything.

Real sessions from our formative study. Every question was typed by a viewer; every answer was generated live.

The story stops when it’s your turn to ask.

0:00

Interaction helped viewers learn

+47.9pts with interaction
+32.1pts without

Pre-to-post learning gain, seven participants who passed the attention test.

Within-participant learning gains

Grey: one participant. Blue: the mean.

Show data

Experience ratings

Four respondents, 1–5 scale.

Show data

06Limitations

Where it still falls short.

Most outputs keep the main context. Perfect scores are rarer.

Score distribution, 72-output analysis sample

Share of outputs per score band.

Show data

Inherited from the video generator

Glass board showing an invented pi-i of tau formula
Invented πi(τ) formulaAWR session · P3
Glass board with the misspelled word Lagrengion and nonsense terms
Misspelled “Lagrengion”AWR session · P5
Benchmark dashboard with plausible-looking but meaningless charts
Charts with no real dataAWR session · P2

Audio drift

Shots are synthesized separately, so voices and ambience can change between them.

Unfaithful symbols

The renderer draws its own text, equations and plots; our verifier checks events, not symbols.

BibTeX

@article{li2026visground,
  title   = {{VIS-Ground}: Video Interactive Storytelling with Contextual Grounding},
  author  = {Li, Bingxuan and Song, Yiwen and Wu, Xueqing and Pan, Yanzhou and Li, Yang and
             Su, Kuang and Liu, Jingyun and Ko, Sebastian and Zhang, Huan and Zhang, Tong and
             Peng, Nanyun and Pfister, Tomas and Song, Yale},
  journal = {arXiv preprint},
  year    = {2026}
}