SessionExpo trackconfirmed

Natively Multimodal from Step Zero

Day
Day 3 — Session Day 2 · Wed, Jul 1
Time
1:55pm-2:15pm
Room
Expo Stage 4 SE
Track
Share
Track theme(approximate)
Agent Architecture & Multimodal Pipelines (Expo)

Sponsor talks on agent architecture decisions (prompt/memory/weights), multimodal models, FDE support agents, and SFT data pipelines.

Accessible with the Expo Explorer pass and above.

About this session

Most AI models start as text systems and have vision, audio, and other modalities added later. That ordering shows up in the work: handoffs between modalities, brittle understanding of mixed inputs, and gaps that surface exactly when real tasks demand reading a chart, a document, and code together. This session looks at a different approach — models trained as multimodal from step zero, where text, image, audio, and video share the same foundation rather than being stitched together. We'll look at why that matters for the kind of work organizations actually want from AI: understanding messy, mixed real-world inputs, holding context across them, and carrying complex tasks end to end. The throughline is what this unlocks for teams deciding where AI can take real work today — and how MiniMax is building toward that frontier.

Speakers

To be announced.