PEGASUS 1.6

media & entertainment

Understand video from a first-person perspective.

A video-to-text model that returns time-coded JSON, now with native image understanding, support for egocentric (first-person) footage, better entity recognition, and cleaner naming.

A video-to-text model that returns time-coded JSON, now with native image understanding, support for egocentric (first-person) footage, better entity recognition, and cleaner naming.

2hrs

Max video duration with full temporal context retained end-to-end.

12

Languages supported across input prompts and generated output.

0

Pre-indexing steps. Send a URL, an asset, or base64, get text back.

JSON

Structured segmentation output ready for your editor, or pipeline.

Things only Pegasus does.

General-purpose VLMs paste frames together and call it understanding. Pegasus is built for video and images, and the gap shows as soon as your prompt mentions a timestamp.

General-purpose VLMs paste frames together and call it understanding. Pegasus is built for video and images, and the gap shows as soon as your prompt mentions a timestamp.

Image

Understands first-person footage.

Bring video descriptions and analysis into the perspective of the person behind the camera with Pegasus 1.6’s expanded understanding of egocentric point-of-view.

Image

Reads the fine print.

On-screen text. Jersey numbers. Whiteboards. Receipts. Pegasus parses the frame alongside the speech track, so summaries include the slides too.

Image

From a raw video to structured JSON.

Define a segment such as a speaker change, brand appearance, or scene cut, choose your fields, and Pegasus returns timestamped JSON.

Image

Sees the frame, not just the script.

Pegasus analyzes images natively, giving it direct access to on-screen text, logos, and visual detail. No bolted-on OCR model, no latency or context-window tax.

Image

The answer comes with a timestamp.

Pegasus answers with exact timestamp, not “somewhere in the middle.” Granular temporal reasoning is built into the model.

Image
Knows who's who.

Improve entity recognition and naming, narrowing the gap with leading models on character/entity ID, which is critical for content ID and character recognition.

From signup to first result in 5 minutes.

Same model, prompts, and JSON output. Choose the surface for your team.

Same model, prompts, and JSON output. Choose the surface for your team.

Python
Node.js
1import requests
2 
3# Step 2: Define the API URL and the specific endpoint
4API_URL = "https://api.twelvelabs.io/v1.3"
5INDEXES_URL = f"{API_URL}/indexes"
6 
7# Step 3: Create the necessary headers for authentication
8headers = {
9 "x-api-key": "<YOUR_API_KEY>"
10}
11 
12# Step 4: Prepare the data payload for your API request
13INDEX_NAME = "<YOUR_INDEX_NAME>"
14data = {
15 "models": [
16 {
17 "model_name": "marengo3.0",
18 "model_options": ["visual", "audio"]
19 }
20 ]
21}

Built for video, now with native image support.

General-purpose multimodal LLMs handle video by sampling stills and creating captions. Pegasus processes the entire video with native image understanding that reads on-screen text, logos, and visual details directly.

General-purpose multimodal LLMs handle video by sampling stills and creating captions. Pegasus processes the entire video with native image understanding that reads on-screen text, logos, and visual details directly.

CAPABILITY

PEGASUS 1.6

Gemini 3.1 PRO

GPT-5.5

Max single-call duration

120 min

90 min

Not specified (omnimodal, no published video duration cap)

Structured segmentation output

JSON-native, schema-conditioned

Structured Outputs supported; no native temporal segmentation

Structured Outputs supported; no native temporal segmentation

Multimodal prompting (image+text)

Yes

Yes

Yes

New domains supported

Ego-centric (first-person) video

TBD

TBD

Start building with TwelveLabs

Bring video intelligence into your production and archive workflows. Deploy, scale, and get more from your content.

Start building with TwelveLabs

Bring video intelligence into your production and archive workflows. Deploy, scale, and get more from your content.

Start building with TwelveLabs

Bring video intelligence into your production and archive workflows. Deploy, scale, and get more from your content.