Claude Agents on Microsoft Foundry: A Setup Walkthrough

The first time I pointed a Claude agent at a Foundry resource, it failed in the most boring way possible: a 401 on a request that looked correct. The endpoint was right, the key was right, and the SDK was one I had used for months. The problem was that I had assumed Foundry's Claude support behaves like the rest of Foundry. It does not, and most of the trouble in enterprise deployments comes from that assumption. Here is the setup path I would follow today, with the gotchas in the order you will hit them.
Claude on Foundry is not the OpenAI-compatible surface
Claude on Foundry uses its own /anthropic endpoint and the Anthropic Messages API, not the OpenAI-compatible endpoint that most Foundry tutorials assume. The base URL is <resource-name>.services.ai.azure.com and Messages calls go to /anthropic/v1/messages` (Microsoft Foundry SDK overview).
This has a practical consequence for anyone with an existing agent codebase. If your orchestration layer speaks Chat Completions, you cannot just swap the model name and move on. You either put a thin adapter in front of the Messages API or you use the Anthropic SDK directly against the Foundry base URL. I prefer the second. The tool-use schema, the content block structure, and the stop reasons all follow Anthropic's format, and hiding that behind an OpenAI-shaped shim means debugging two abstractions when something goes wrong.
There is also a naming trap. In inference calls, the model parameter takes the deployment name you chose in Foundry, not necessarily the model ID (Microsoft Foundry Claude guide). If you named your deployment something friendly, use that string. I now keep deployment names identical to model IDs so there is one fewer mapping to get wrong in config.
Auth: three things that produce a 401
Most Foundry-Claude 401s come from one of three causes: wrong token scope, a missing role assignment, or an expired token in a long-running process. I have seen all three in a single week on one project.
Token scope. For Entra ID authentication, Microsoft's troubleshooting table says the scope must be ai.azure.com and a wrong scope is a listed cause of 401 ([Microsoft Foundry Claude guide](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/how-to/use-foundry-models-claude)). Older third-party tutorials use cognitiveservices.azure.com If you copied a snippet from a blog post written a while ago, check this first.
Role assignment. Calling the Messages API with Entra ID requires the Cognitive Services User role on the Foundry account (Claude on Foundry starter kit). My own recommendation is to assign that role explicitly to the managed identity or service principal your agent runs as, rather than relying on broader roles you may already have on the subscription. It is one extra line of infrastructure code and it removes a whole class of ambiguity when you debug a 401.
Token lifetime. Entra ID tokens for Claude on Foundry typically expire after 1 hour (Claude Platform docs). A scheduled worker that grabs a token at startup and reuses it will work in testing and die at the 61-minute mark in production. Use a token provider that refreshes, not a static token string.
Here is the shape I use in Python with the Anthropic SDK and azure-identity:
from anthropic import AnthropicFoundry
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
token_provider = get_bearer_token_provider(
DefaultAzureCredential(),
"ai.azure.com
)
client = AnthropicFoundry(
base_url="<resource-name>.services.ai.azure.com
azure_ad_token_provider=token_provider,
)
Check the current Anthropic SDK docs for the exact client class and parameter names for your SDK version, since these have shifted between releases. The important part is the pattern: the provider is a callable that refreshes, and the scope is the official one. For local development, DefaultAzureCredential picks up your az login session. In production it should resolve to a managed identity, which also removes API keys from your environment entirely, something every enterprise security review asks about.
Choosing a hosting version and a model
Foundry offers Claude in two hosting versions: Hosted on Azure and Hosted on Anthropic infrastructure. All Claude models support Global Standard, while Data Zone Standard (US) is available only for some Hosted on Azure models, including claude-haiku-5-5, claude-opus-5-5, claude-opus-5, claude-opus-4-8, claude-sonnet-5, claude-sonnet-5-5 and claude-haiku-4-5 (Claude models in Foundry).
If a client has data residency requirements, this is the first conversation to have, before any code. Global Standard is the simplest deployment type, but it may not satisfy a compliance team that wants processing kept inside a defined zone. Data Zone Standard (US) does, for the models listed. Which hosting version and deployment type you pick is a legal and architectural decision, so settle it early.
On model choice, my split in multi-agent systems is consistent: a stronger model for planning and judgment steps, a cheaper and faster one for high-volume, well-bounded steps like extraction, classification and routing. On Foundry, claude-opus-5-5, claude-sonnet-5-5 and claude-haiku-5-5 have a 1M-token context window and 128K max output, while claude-haiku-4-5 has 200K context and 64K max output (Claude models in Foundry).
Two warnings about model versions:
- Pin explicit versions. Anthropic's docs say to pin an explicit model version in Azure deployments instead of choosing auto-update to latest (Claude Code on Foundry). An agent whose behavior changes under you because a deployment auto-upgraded is a nightmare to debug.
- Do not assume a new model is a drop-in. The Claude Platform release notes for Haiku 5.5 (listed under October 7, 2026) explicitly warn that code written for Haiku 4.5 can break on Haiku 5.5 (release notes). Run your eval set before you switch a tier.
Deprecations matter here too. Anthropic announced on Sept 30, 2026 that Claude Sonnet 4.5 (claude-sonnet-4-5-20250929) is deprecated, with Claude API retirement scheduled for Nov 30, 2026, and recommends migrating to Sonnet 5.5 (release notes). If any of your agents still pin Sonnet 4.5, put the migration on the calendar now. I have not verified whether the same dates apply to the Foundry deployment of that model, so check the Foundry model page before you plan around them.
Tool use and orchestration
The tool loop itself is the standard Messages API pattern: send tools, receive tool_use blocks, execute them, return tool_result blocks, repeat until the model stops. Nothing about Foundry changes that. What changes is the surrounding plumbing and one model-specific breaking change.
For Claude Opus 5.5, thinking cannot be disabled and forced tool use returns an error (Opus 5.5 docs). If any of your agents relied on forcing a specific tool call to get deterministic routing, that pattern breaks on Opus 5.5. The fix is to describe the routing requirement in the system prompt and the tool descriptions, and to validate the chosen tool in your own code rather than relying on the API to force it. I would rather have that validation anyway: a router that trusts the model blindly is a router that will eventually call the wrong tool.
For orchestration you have a few options. Microsoft's Claude documentation lists the Microsoft Agent Framework and the Claude Agent SDK as ways to build agents with Claude on Foundry (Claude models in Foundry), and Microsoft's GA announcement says Foundry Agent Service orchestrates multi-step agents that use Claude as their reasoning core (Microsoft blog).
Here is where I have to be careful, because I have not tested one thing. It is unclear whether Claude works with Foundry Agent Service's managed create_agent (threads/runs) API with portal visibility. Sources I found disagree on whether it is supported. Test this against your own resource before you design around the managed agent runtime. If it does not work for you, the fallback is what I use anyway: run the agent loop in your own code, call Claude through the Messages API, and treat Foundry as the model host. You lose portal-native agent traces but gain full control over state, retries and tool execution, and you can wire your own tracing in.
My orchestration pattern for multi-agent work is simple: one coordinator agent that plans and delegates, specialist agents with narrow tool sets, and a shared state object that the coordinator owns. Specialists never call each other directly. This keeps failures local and makes each agent's prompt and tool list small enough to evaluate on its own.
Quotas and latency: where enterprise deployments get surprised
Default quotas differ sharply by subscription type, and the gap is large enough to decide whether your architecture works at all.
| Subscription type | Models | RPM | ITPM | OTPM |
|---|---|---|---|---|
| Pay-as-you-go | claude-opus-5-5, claude-sonnet-5-5, claude-sonnet-5 | 40 | 40,000 | 8,000 |
| Pay-as-you-go | claude-haiku-4-5, claude-sonnet-4-6, claude-sonnet-4-5 | 80 | 80,000 | 16,000 |
| Enterprise / MCA-E | claude-opus-5-5, claude-sonnet-5-5, claude-haiku-4-5 | 10,000 | 10,000,000 | 2,000,000 |
Source: Claude model quotas and rate limits. Quotas change over time, so confirm in the portal. Haiku 5.5 limits were not in the table I could verify.
Do the arithmetic on the pay-as-you-go row before you promise anything. At 40,000 input tokens per minute, a single agent turn that carries a 20,000-token context uses half the minute's budget. A multi-agent run with several parallel specialists will hit 429s almost immediately. The enterprise defaults are a different world, so find out which subscription type your client's deployment sits on during discovery, not during load testing. There is conflicting guidance on eligibility: some sources say Enterprise or MCA-E subscriptions are required, while the quota page lists pay-as-you-go defaults. Availability depends on subscription, so verify in the portal.
Microsoft's guidance for rate limiting is to handle HTTP 429 with exponential backoff (Claude models in Foundry). In practice I add jitter, cap total retry time, and put a concurrency limiter in front of the client so the agent queues work instead of firing a burst and then backing off in unison.
On latency, I have no official Foundry figures or SLAs to give you. Vendor speed claims for specific models are not measurements of your deployment, so do not plan capacity from them. Measure time to first token and total turn time per model tier on your own resource, in your own region, with your own prompts, and keep those numbers in your eval harness so you catch regressions when a model version changes. Region availability is something I also cannot confirm from an official page, so check the Foundry portal for the regions your deployment type supports.
A verification checklist
Before calling a Foundry-Claude agent production-ready, I run through this:
- A call with an intentionally wrong scope returns 401, and the correct scope succeeds (proves you know which failure is which).
- The identity has the Cognitive Services User role, verified with a managed identity and not your personal login.
- A process left running past one hour still authenticates (token refresh works).
- The
modelfield matches the deployment name, and the deployment pins an explicit version. - A synthetic burst above your RPM limit produces 429s that your backoff handles without dropping work.
- The tool loop is tested with a model swap, since forced tool use fails on Opus 5.5 and Haiku 4.5 code can break on Haiku 5.5.
- Cost is measured from real token counts. Anthropic's API list prices are not confirmed Foundry prices, so read your own Azure billing.
What I'd do
Start with the Anthropic SDK against the /anthropic endpoint and a managed identity. Skip the abstraction layers until you have a working loop and an eval set. Pin versions, and treat every model upgrade as a migration with tests, not a config change.
Find out the subscription type and the data residency requirement in the first meeting. Those two facts decide your quota ceiling and your deployment type, and both are painful to change late.
Keep orchestration in your own code unless you have confirmed that the managed agent runtime supports Claude the way you need. Foundry as model host, your code as the control plane, is the setup I trust to run unattended.
If you are wiring Claude into Foundry for something with real stakes and want a second pair of eyes on the architecture, you can reach me at lazar-milicevic.com/#contact. Otherwise, the rest of the blog covers the production patterns behind the agent systems I build.
Frequently asked questions
Is Claude on Microsoft Foundry compatible with the OpenAI Chat Completions API?
No. Claude on Foundry uses its own /anthropic endpoint and the Anthropic Messages API, not the OpenAI-compatible endpoint most Foundry tutorials assume. Calls go to /anthropic/v1/messages on your resource's services.ai.azure.com base URL. If your orchestration layer speaks Chat Completions, you need a thin adapter or you can use the Anthropic SDK directly. I prefer the SDK route because tool-use schemas, content blocks and stop reasons all follow Anthropic's format, and an OpenAI-shaped shim means debugging two abstractions.
Why am I getting a 401 error when calling Claude on Foundry with Entra ID?
In my experience there are three common causes: the wrong token scope, a missing role assignment, or an expired token. The scope should be ai.azure.com, not the older cognitiveservices.azure.com that many tutorials use. The identity also needs the Cognitive Services User role on the Foundry account, which I recommend assigning explicitly to your managed identity or service principal. Finally, Entra ID tokens typically expire after about an hour, so long-running workers need a refreshing token provider rather than a static token.
What value should I use for the model parameter when calling Claude on Foundry?
Use the deployment name you chose in Foundry, which is not necessarily the model ID. If you gave your deployment a friendly custom name, that exact string goes in the model parameter. To avoid a mapping mistake in config, I keep deployment names identical to the model IDs. That leaves one fewer thing to get wrong when you move between environments.
How do I handle data residency requirements for Claude on Microsoft Foundry?
Settle it before writing any code, because it is a legal and architectural decision. Foundry offers Claude in two hosting versions, Hosted on Azure and Hosted on Anthropic infrastructure, and all Claude models support Global Standard deployment. Data Zone Standard (US) is available only for some Hosted on Azure models, such as claude-opus-5-5, claude-sonnet-5-5 and claude-haiku-4-5. If a compliance team wants processing kept inside a defined zone, Global Standard may not be enough, so check the model list first.
How should I choose between Claude models for a multi-agent system on Foundry?
I use a consistent split: a stronger model for planning and judgment steps, and a cheaper, faster one for high-volume, well-bounded steps like extraction, classification and routing. Context limits also matter. On Foundry, claude-opus-5-5, claude-sonnet-5-5 and claude-haiku-5-5 offer a 1M-token context window and 128K max output, while claude-haiku-4-5 offers 200K context and 64K max output. Match the model to the step so you are not paying for capability a routing task never uses.
Building something hard with AI or automation? I am open to talk.
Get in touch