Understand video from a first-person perspective.
2hrs
Max video duration with full temporal context retained end-to-end.
12
Languages supported across input prompts and generated output.
0
Pre-indexing steps. Send a URL, an asset, or base64, get text back.
JSON
Structured segmentation output ready for your editor, or pipeline.
Things only Pegasus does.

Understands first-person footage.
Bring video descriptions and analysis into the perspective of the person behind the camera with Pegasus 1.6’s expanded understanding of egocentric point-of-view.

Reads the fine print.
On-screen text. Jersey numbers. Whiteboards. Receipts. Pegasus parses the frame alongside the speech track, so summaries include the slides too.

From a raw video to structured JSON.
Define a segment such as a speaker change, brand appearance, or scene cut, choose your fields, and Pegasus returns timestamped JSON.

Sees the frame, not just the script.
Pegasus analyzes images natively, giving it direct access to on-screen text, logos, and visual detail. No bolted-on OCR model, no latency or context-window tax.

The answer comes with a timestamp.
Pegasus answers with exact timestamp, not “somewhere in the middle.” Granular temporal reasoning is built into the model.

Knows who's who.
Improve entity recognition and naming, narrowing the gap with leading models on character/entity ID, which is critical for content ID and character recognition.
From signup to first result in 5 minutes.
Built for video, now with native image support.
CAPABILITY
PEGASUS 1.6
Gemini 3.1 PRO
GPT-5.5
Max single-call duration
120 min
90 min
Not specified (omnimodal, no published video duration cap)
Structured segmentation output
JSON-native, schema-conditioned
Structured Outputs supported; no native temporal segmentation
Structured Outputs supported; no native temporal segmentation
Multimodal prompting (image+text)
Yes
Yes
Yes
New domains supported
Ego-centric (first-person) video
TBD
TBD


