System Design for AI Systems: Caching, Rate Limits & State Persistence
Architectural blueprints for scaling agentic systems to production: semantic caching, backpressure handling, distributed state, and failure recovery.
On this page
Deploying an autonomous agent or multi-step LLM workflow in production presents challenges that traditional web services never face: unpredictable execution latency (from 2 seconds to 45 seconds), volatile API rate limits, non-deterministic cost profiles, and stateful multi-turn interactions.
Architectural Architecture: Decoupled Job Queues
Never execute long-running agent workflows synchronously within an HTTP request/response cycle. If the client disconnects or an upstream model hangs, resources are orphaned.
Instead, implement an asynchronous job queue:
[Client] ---> POST /api/tasks ---> [API Gateway] ---> [Redis / SQS]
|
v
[Client] <--- SSE / Polling <--- [State DB] <--- [Agent Worker Pool]Semantic Caching with Vector Thresholds
For high-volume query workloads, identical or semantically identical prompts should bypass the LLM entirely:
import numpy as np
def cosine_similarity(v1, v2):
return np.dot(v1, v2) / (np.linalg.norm(v1) * np.linalg.norm(v2))
def check_semantic_cache(query_vector, cache_records, threshold: float = 0.95):
"""Return cached answer if query is semantically indistinguishable."""
for record in cache_records:
sim = cosine_similarity(query_vector, record["vector"])
if sim >= threshold:
return record["cached_response"]
return NoneImplementing Token Bucket Rate Limiting
To avoid sudden 429 Too Many Requests errors from model providers, implement client-side token bucket limiters that buffer spikes and enforce smooth request cadence.
By treating AI agents as distributed, asynchronous systems rather than simple HTTP handlers, your systems maintain high availability even under extreme upstream latency and provider volatility.
Keep learning with Sri
More practical tutorials and experiments on the channel.