Squiz, a global Digital Experience Platform provider, is transforming how organizations deliver conversational search experiences. By adopting Amazon S3 Vectors, Squiz reimagined its ingestion pipeline — increasing data processing speed by 50% and shifting from bespoke, always-on infrastructure to a scalable serverless model. This allows Squiz to seamlessly scale from 25,000 to millions of vectors per client, while significantly reducing costs. Hear how this shift freed engineering teams to focus on RAG innovation rather than infrastructure management, and how it powers smart video search capabilities across their platform.
What this session is about
Live updates related to this session LIVE
Sourced via Parallel AI Monitor — continuous web watch on 21 topical streams. Updated .
- truefoundry.com Scaling infra for agent workloads
AI Platform Engineering: A Complete Guide for 2026
Refonte Learning published an API design guide for AI agents recommending rate limits keyed to the authenticated principal and adjusted for operation cost, rather than relying only on shared infrastructure limits. This provides concrete guidance for managing tool-call and API loa
- oodaloop.com Scaling infra for agent workloads
Scaling AI agents with trustworthy data - OODAloop
Gravitee published an update explaining that the stateless Model Context Protocol (MCP) specification simplifies horizontal scaling for agent workloads. The update also highlights AI gateway governance as a way to manage agent tool traffic, security, and scaling across MCP-based
- docs.cloud.google.com high confidence Scaling infra for agent workloads
Scale your agents | Gemini Enterprise Agent Platform | Google Cloud Documentation
Waxell published 'AI Agent Cost Enforcement: Before vs. After [2026]' on June 24, 2026, outlining a shift from post-execution cost visibility to pre-execution hard enforcement. This architectural change allows for per-task budget ceilings and the immediate termination of runaway
- medium.com Agent memory & RAG architectures
Building Memory for AI Agents: Context Windows, Vector ...
Redis published a 2026 State of Context Engineering report based on a survey of IT and AI infrastructure leaders. It frames production agent context as a combination of data, records, memory, and live state, and describes context engineering as selecting the facts, past interacti
- parallel.ai Scaling infra for agent workloads
Parallel - Web Infrastructure for AI Agents
Refonte Learning published an API design guide for AI agents recommending rate limits keyed to the authenticated principal and adjusted for operation cost, rather than relying only on shared infrastructure limits. This provides concrete guidance for managing tool-call and API loa
External links matched to this session via topic relevance. The KB does not endorse third-party content; verify before citing.