Bedrock and Vertex together: a hybrid AI platform built to survive losing either one
two providers, one gateway, and a lock-in budget we actually drill. agent work on bedrock, warehouse retrieval on vertex, and a path for teams to onboard without ever importing a provider SDK.
- no dominant cloud deal, no pre-approved vendor lane, no identity plane that forced AWS or GCP. the decision got made on workload fit.
- agents and tool-calling run on bedrock (AgentCore, Strands). warehouse-adjacent retrieval runs on vertex (BigQuery). each provider does the job it is already good at.
- every call goes through one LiteLLM gateway: routing, virtual keys, caching, budgets, circuit breakers per provider-and-model pair. application code never sees a Bedrock or Vertex SDK.
- swapping an inference call is a config change. swapping a fine-tuned model is a retrain. we track that as a lock-in budget and prove it with a provider-swap gameday.
- observability is three layers: D.U.R.E.S.S. for the gateway, OpenTelemetry GenAI conventions for the model calls, a separate quality and safety layer for the canary gate.
- the workloads that win: on-call agents, warehouse RAG, mixed retrieve-then-act copilots, regulated intake. the ones that lose: anyone who bypasses the gateway "just this once."
Why procurement didn't pick the model
Going in we had no enterprise agreement pinning us to one cloud, no single pre-approved AI vendor, and no identity plane that made AWS or GCP the default. Plenty of orgs treat that as unfinished homework. It was the opening. Agent runtime and data gravity went on the table first. Cloud commitment followed the workload.
The question that actually mattered: what happens the day one of these vendors stops being the right choice. Two providers means two MSAs, two DPAs, two FinOps forecasts, committed-use on both sides, and a failover you can actually drill. Warehouse retrieval stays next to BigQuery. Guardrails, audit trail, and virtual keys sit on the gateway. App code does not start with boto3. The rest of this write-up is the architecture that makes that answerable in hours, not a quarter.
The decision framework
Six factors, in the order they actually carried weight. The first three came back flexible, which is what let the last three get decided on merit.
| factor | our answer | effect on the decision |
|---|---|---|
| cloud commitment | flexible, no dominant enterprise agreement | didn't force a platform on its own |
| compliance and procurement | flexible, no single pre-approved vendor boundary | same |
| identity plane | flexible | same |
| model and agent tooling | bedrock: broadest single-API model swap, most mature agent runtime (AgentCore, Strands) | owns the agent-workload lane |
| data gravity | vertex: BigQuery-native retrieval, matches an existing data-heavy workload profile | owns the data-heavy lane |
| lock-in cost | fine-tuning deltas don't move between platforms, gateway abstraction only insulates inference calls | its own section, further down |
Why this matters at a system level
Agent workloads, on-call routing, anything doing multi-step tool calling, run on Bedrock. Retrieval and analysis that already sits next to the warehouse run on Vertex. That split is the daily operating model. The reasons below are why we paid for a second control plane.
Blast radius. One vendor means a provider outage, a rate-limit change, a silent model version bump, or a content-filter update degrades every AI-dependent system at once. Split by workload type and a Bedrock incident degrades agents while warehouse retrieval on Vertex keeps running, and the reverse. Same reasoning that keeps production out of a single availability zone, one layer up the stack.
Roadmap independence. A single vendor's roadmap becomes the platform's capability ceiling. AgentCore matured faster on Bedrock than Vertex's agent tooling did. BigQuery-native retrieval matured faster on Vertex than Bedrock Knowledge Bases did. One vendor means inheriting whichever of those two timelines that vendor is behind on.
Negotiating position. Meaningful spend on two hyperscalers is a different conversation than nominal multi-cloud with almost all of the money on one side. Committed-use structures differ across AWS and GCP. Real leverage on both sides of a renewal is worth more than the extra operational surface costs.
Shadow AI. If the governed platform cannot serve a workload well, teams route around it: outside the gateway, outside the audit trail, outside every control this platform exists to provide. A bad fit for a real chunk of the work does not stay a clean architecture. It becomes the thing engineers quietly bypass.
None of this is an argument for a third or fourth provider. Operational surface grows with every additional platform, and the return drops fast past two. Two providers, chosen because they are genuinely differentiated on the axes that matter, is where the resilience and leverage still outweigh the cost of a hybrid control plane.
The architecture
Every call from an internal service goes through the same path, regardless of which provider ends up handling it.
The gateway is the operating discipline. Four things wrap every request, and they live here so they do not get reimplemented per service:
model_list:
- model_name: agent-tasks
litellm_params:
model: bedrock/us.anthropic.claude-sonnet-4-5-20250929-v1:0
aws_region_name: us-east-1
- model_name: data-heavy-tasks
litellm_params:
model: vertex_ai/gemini-2.5-pro
vertex_project: your-project-id
vertex_location: us-central1
router_settings:
routing_strategy: usage-based-routing-v2
fallbacks:
- agent-tasks: [data-heavy-tasks]
cooldown_time: 30
litellm_settings:
cache: true
cache_params:
type: redis
ttl: 600
budget_manager: true
How a team actually gets onto Bedrock
The hybrid diagram is the easy part. The failure mode is a team that "just needs Claude for this one service" and ships a boto3 client with an IAM user key in a secret. That call never hits the gateway, never hits a budget, never hits the canary, and never fails over.
- Platform owns model access and IAM. Enable the model in the Bedrock console once. The gateway's role holds
bedrock:InvokeModelandbedrock:InvokeModelWithResponseStream. Prefer a VPC endpoint so tokens never leave the private network. Application roles do not get Bedrock permissions. - Alias the model, do not leak the ID. App code calls
agent-tasks, notus.anthropic.claude-sonnet-4-5-20250929-v1:0. The gateway maps the alias to a Bedrock inference profile. Theus.prefix is AWS cross-region inference, useful as a same-cloud failover. It is not a second provider. - Issue a virtual key and a budget. Per team, per environment. LiteLLM throttles at the gateway when the budget burns. A billing alert a week later is how you find out someone looped an agent overnight.
- Point the existing OpenAI-compatible client at the gateway. Base URL plus the virtual key. No Bedrock SDK, no Vertex SDK, no AWS keys in the app. If the team already speaks OpenAI, the code change is the endpoint.
- Strands goes through the gateway. AgentCore just hosts it. AgentCore is the deployment runtime. Strands is the agent framework that runs on it. The piece that actually routes inference is Strands'
LiteLLMModel, imported fromstrands.models.litellm. Gateway URL and virtual key go inclient_args.model_idis the gateway alias. ABedrockModel(or aLiteLLMModelthat never enters proxy mode) calls AWS directly and rebuilds the bypass. - Promote in that order: one model, logging only, no fallbacks. Then a Vertex fallback on the same alias. Then Redis caching for the cacheable prompts. Then the canary gate on prompt and model version changes. Do not turn all four on in the first PR.
from strands import Agent
from strands.models.litellm import LiteLLMModel
model = LiteLLMModel(
client_args={
"api_base": "https://your-gateway/v1",
"api_key": "<virtual-key>",
"use_litellm_proxy": True,
},
model_id="agent-tasks",
params={"max_tokens": 1000},
)
agent = Agent(model=model, tools=[...])
use_litellm_proxy: True (or a litellm_proxy/ prefix on model_id) is what makes LiteLLM treat api_base as your gateway. Without it, LiteLLM may still call Bedrock directly. Spend, fallbacks, caching, and the canary gate stay in LiteLLM either way. AgentCore is not in that path.
Stand up a Bedrock Knowledge Base for data that already lives in BigQuery. You will copy the warehouse into OpenSearch and then wonder why Vertex was in the design.
Fine-tune on Bedrock because a prompt "wasn't good enough yet." Fine-tuned weights do not move. RAG and prompt work are the default until a measured quality gap says otherwise.
Run legacy Bedrock Agents, AgentCore, and a homegrown graph in the same namespace. Pick one runtime. Wire Strands through LiteLLMModel either way.
Guardrails stay on both sides. Bedrock Guardrails for PII, topic deny, and prompt-attack filters on the AWS path. Gateway-level logging and redaction for every call, including Vertex. A Bedrock-only guardrail does nothing for the warehouse RAG route.
Workloads that actually win
A hybrid control plane is wasted on a single chatbot with no tools and no warehouse. These are the shapes that pay for the extra surface.
On-call and incident agents
Runbook lookup, ticket updates, paging, Kubernetes reads. Long tool chains, tight latency, needs a mature agent runtime. Fallback to Vertex keeps the agent answering when Bedrock quota burns during an incident, which is when you least want a 429.
Warehouse RAG
Claims, finance close, ops metrics, anything whose source of truth is already in BigQuery. Do not ETL that into a Bedrock Knowledge Base so the agent stack can pretend it is one cloud. Retrieval stays next to the data.
Retrieve, then act
Internal copilots that pull a number from the warehouse and then open a ticket, draft a change, or call a tool. Two model aliases, same OpenAI-compatible client. This is the workload the gateway exists for.
Regulated intake
Documents with PHI, PII, or payment data. Redact at the gateway, apply Bedrock Guardrails on the AWS path, keep an audit trail per virtual key. Cheapest-model routing is how this leaks.
High-volume triage
Ticket classification, change-advisory summaries, FAQ deflection. Cacheable, bursty, sensitive to unit cost. Model tiering and the Redis cache are where the 25 to 35 percent cost drop actually comes from.
Prompt and model promotions
A second provider makes "incumbent vs candidate" a routing problem, not a rewrite. Shadow traffic scores quality, latency, and cost before a silent provider bump ships into production.
A team with one model, no tools, no warehouse, and no compliance boundary. Give them a single alias and a budget. Add the second provider when a real workload needs it, not because the architecture diagram has two boxes.
The portability problem
Two providers means two ways to get locked in, and they do not share a fix.
The gateway boundary covers inference. No application code calls the Bedrock or Vertex SDK. Everything goes through LiteLLM's OpenAI-compatible interface. A provider swap at the call level is a config change at the gateway, not a rewrite across services. That is most of what people mean when they say "avoid lock-in."
Fine-tuned weights are a different problem. A model fine-tuned on Bedrock has to be retrained from base weights to run on Vertex. Neither vendor exports the fine-tuning delta in a format the other can read. The gateway ends at the inference call. It has nothing to say about weights sitting on one side of the split.
We track vendor lock-in the way we'd track an error budget. Two rules: default to RAG and prompt work over fine-tuning wherever the quality bar allows it, so the dependency is avoided rather than managed. Where fine-tuning is required, the swap cost gets estimated and written down before the model ships, not discovered during an outage or a contract renegotiation. That estimate is a number we can be wrong about and correct.
We keep that number honest with a periodic gameday that executes a provider failover. Production-representative traffic reroutes away from one provider through the gateway only, no code deploy. We measure how much of the request volume keeps working and how long full rerouting takes. A fine-tuned model sitting in the failover path surfaces immediately, because the gateway can reroute the inference call but cannot reroute the weights.
Observability: three layers
D.U.R.E.S.S. is the platform layer. It answers whether the gateway itself is healthy, the same question it answers for any other piece of infrastructure. It does not answer whether a model's output was any good, which step in a multi-tool agent run failed, or whether a response leaked something it shouldn't have. Treating D.U.R.E.S.S. as the whole story is how you walk into an incident with green dashboards and a bad model.
Layer 1, platform signals
| signal | classic meaning | AI gateway meaning |
|---|---|---|
| Duration | request latency | p50 and p95 token latency per model |
| Utilization | resource usage vs capacity | token throughput against provider quota |
| Rate | requests per second | requests per second per model route |
| Errors | 4xx and 5xx | 4xx, 5xx, and circuit breaker trips |
| Saturation | proximity to limits | proximity to provider rate limits |
| System health | uptime, dependency health | circuit breaker state per provider-model pair |
Layer 2, GenAI telemetry
Instrumented against the OpenTelemetry GenAI semantic conventions rather than a bespoke schema, specifically so it holds up across both providers without a rewrite. Captured per call: requested and actually-served model plus version, input, output, and reasoning token counts, time to first token and tokens per second rather than one blended duration number, and finish reason (stop, length, tool call, content filter). For agent workloads, every tool call, model invocation, and retrieval step becomes a child span, producing a full trace of the reasoning chain instead of a single top-level result.
Layer 3, quality and safety
The layer that feeds the canary gate: groundedness and hallucination scoring against retrieved context, guardrail trigger events (PII redaction, prompt injection detection, content filter hits), and output safety scores. Kept separate from layer 2 on purpose. Observability is the lossless capture of what happened: a span, a token count, a tool call argument. Evaluation is scoring whether what happened was good. Mix them and a wrong scorer silently corrupts the record of what occurred.
The canary gate
Model and prompt version changes are deployments. A new version runs against shadow traffic before it promotes, scored on output quality, latency, and cost delta against the incumbent. The pattern is borrowed from a Kayenta-based canary platform built at a prior company, which ran over 10,000 analyses a quarter and blocked roughly 11 percent of deploys before they shipped. That block rate is the reference the gate's thresholds were calibrated against.
A silent model version bump from a provider is the same class of risk as a bad deploy from your own team. The gate runs the new version against shadow traffic, scores it, and blocks promotion if the numbers regress. The provider does not get a free pass because the change came from their side.
Implementation notes
Gateway
LiteLLM runs as its own deployment behind the existing ingress rather than a sidecar per service. Virtual keys and budgets are issued per team, and circuit breakers are scoped per provider-and-model pair specifically, so a Bedrock outage on one model doesn't take down a working Vertex fallback route.
FinOps
Cost data from both providers is normalized through a FOCUS-aligned pipeline, the FinOps Open Cost and Usage Specification, so spend is reported on a single schema regardless of which vendor generated it. Per-team budgets are enforced at the gateway through LiteLLM's budget manager, not reconciled after the fact. A team hitting its budget gets throttled at the gateway, not surprised by a billing alert a week later.
The scoreboard
The hybrid split shows up clearest in four numbers.
The cost reduction is a range, not a single figure, because it moves with traffic mix. A quarter heavy on cacheable agent tasks trends toward the top of it; a quarter heavy on uncached data-heavy retrieval trends toward the bottom. Reporting it as one number would be cleaner and less honest.
The canary block rate lines up with the 11 percent reference from the prior Kayenta platform, which is the point. A block rate far below that would mean the gate is too loose to catch real regressions. A block rate far above it would mean the gate is throttling legitimate changes. Sitting in the same band as a proven platform is where it should be.
Beyond those four, the framework should always be reporting these, whether it is this platform or the next one:
- per-route token throughput against provider quota, the Utilization signal that catches a saturation problem before it becomes a latency problem
- circuit breaker trips per provider-model pair, since a trip on one route should not cascade into a trip on an unrelated fallback route
- gameday failover success rate and reroute time, per drill, because the lock-in budget is only as honest as the last time it was actually tested
- fine-tuned models in the failover path, as a count, since the gateway can reroute the call but not the weights, and a growing count is the early warning that the lock-in budget is being spent
- shadow traffic quality delta per pending promotion, the number that actually decides whether a model or prompt version ships, not a fixed timer
What we would tell the next team doing this
- 01Make the governed path the fastest path. Virtual key, alias, OpenAI-compatible client. If boto3 is easier, teams will use boto3, and the gateway becomes a diagram.
- 02AgentCore hosts. Strands agents. LiteLLMModel infers. Point
LiteLLMModelat the gateway withclient_argsand the alias asmodel_id. ABedrockModel, or a LiteLLM client that never enters proxy mode, skips budgets, fallbacks, and the canary. - 03Swapping an inference call is a config change. Swapping a fine-tuned model is a retrain. Default to RAG and prompt work. Where fine-tuning is required, write the swap cost down before the model ships.
- 04Run a real provider-swap gameday. Production-representative traffic, gateway-only reroute, no code deploy. A fine-tuned model in the failover path will surface in minutes.
- 05Keep observability and evaluation in separate layers. Observability is the lossless record of what happened. Evaluation is an opinion about whether it was good. Mix them and a wrong scorer corrupts the record.
- 06Two providers is where the math still works. Add a third when a workload neither provider serves well. A team with one model and no warehouse does not need the hybrid diagram on day one.
- 07Treat a provider's silent model version bump as your own bad deploy. Shadow traffic, score it, block it. The provider does not get a free pass because the change came from their side of the API.