
Calling the Claude API from Python: A Practical Guide
Adding a large language model to a Python app usually starts with a five-line script that works on the first try. Then real requirements show up: responses that need to stream into a UI, output that has to be valid JSON, the model calling your own functions, a long document sent with every request, rate limits, and timeouts. Each of those has a right way to do it, and the official SDK already handles most of the plumbing if you know where to look.
This guide walks through calling Anthropic's Claude models from Python with the official anthropic SDK. It covers setup and your first request, choosing a model, conversations and system prompts, streaming, effort and thinking, structured output with Pydantic, tool use, prompt caching, async requests, and error handling. Every example uses the current Messages API and model IDs as of October 2026.
Setup
Install the SDK into a virtual environment:
python -m pip install anthropic
The 1.x series of the SDK requires Python 3.10 or newer. If you're upgrading an older project from 0.x, note that 1.x is built on httpx2 (a maintained fork of httpx) rather than httpx, and long-deprecated APIs like Text Completions were removed. The SDK's migration guide on GitHub covers the details.
Create an API key in the Claude Console and export it as an environment variable:
export ANTHROPIC_API_KEY="sk-ant-..."
The client reads ANTHROPIC_API_KEY automatically, so you never need to put the key in your code. In production, load it from your secrets manager or deployment environment, and keep it out of version control.
Your First Request
# first_call.py
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
messages=[
{"role": "user", "content": "Explain Python's GIL in two sentences."},
],
)
for block in response.content:
if block.type == "text":
print(block.text)
print(response.stop_reason)
print(response.usage.input_tokens, response.usage.output_tokens)
Everything goes through one endpoint, the Messages API, via client.messages.create(). Three arguments are required:
model: the model ID.claude-opus-5-5is the current Opus model.max_tokens: the maximum number of tokens the model may generate. It's a hard cutoff, not a target length, so don't set it too low. If the response hits the limit, it's truncated mid-sentence.messages: the conversation so far, as a list of{"role": ..., "content": ...}dicts. The first message must be from theuser.
The response is a Message object. Its content is a list of content blocks, not a single string. A plain answer is usually one text block, but responses can also include thinking blocks and tool_use blocks, so always check block.type before reading block.text. Code that does response.content[0].text works until the day the first block is something else.
stop_reason tells you why the model stopped:
stop_reason | Meaning |
|---|---|
end_turn | The model finished normally |
max_tokens | Hit your max_tokens limit; output is truncated |
stop_sequence | Hit one of your stop_sequences |
tool_use | The model wants you to run a tool |
pause_turn | A long server-side tool turn paused; send it back to continue |
refusal | The model declined the request |
usage reports input and output tokens, which is what you're billed on. Log it from day one; it's the basis for any cost analysis later.
Choosing a Model
Three current models cover most needs:
| Model | ID | Context window | Input / output price per 1M tokens |
|---|---|---|---|
| Claude Opus 5.5 | claude-opus-5-5 | 1M tokens | $4 / $20 |
| Claude Sonnet 5.5 | claude-sonnet-5-5 | 1M tokens | $2 / $10 |
| Claude Haiku 4.5 | claude-haiku-4-5 | 200K tokens | $1 / $5 |
Opus is the most capable of the three and a sensible default for complex reasoning, coding, and agentic work. Sonnet balances capability and cost for high-volume production use. Haiku is the fastest and cheapest, a good fit for classification, routing, and simple extraction. Prices and model lineups change, so check the models overview before committing to numbers in a budget. You can also query the Models API from code with client.models.list() and client.models.retrieve("claude-opus-5-5"), which return each model's context window and capabilities.
Use the model ID exactly as shown. Don't append date suffixes you might have seen for older models.
System Prompts and Conversations
A system prompt sets the role, rules, and context for the whole conversation. Pass it with system=, separate from messages.
The API is stateless: it doesn't remember previous calls. For a multi-turn conversation, you send the full history every time and append the assistant's reply before the next user message:
# chat.py
import anthropic
from anthropic.types import MessageParam
client = anthropic.Anthropic()
SYSTEM = "You are a concise Python tutor. Prefer short answers with one code example."
history: list[MessageParam] = []
def ask(question: str) -> str:
history.append({"role": "user", "content": question})
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system=SYSTEM,
messages=history,
)
history.append({"role": "assistant", "content": response.content})
return "".join(block.text for block in response.content if block.type == "text")
print(ask("What does enumerate() do?"))
print(ask("Show the same thing with a start index of 1."))
The second question only makes sense because the first exchange is in history. Two details matter here:
- Append
response.content, not just the extracted text. The content blocks may include thinking blocks and tool calls that need to go back to the API unchanged in later turns. Passing the full list keeps the history valid. - History grows with every turn, and you pay for those input tokens each time. For long-running conversations, prompt caching (covered below) cuts that cost substantially.
MessageParam is the SDK's type for a message dict. Use the SDK's types rather than defining your own, and a type checker will catch malformed messages before the API does.
One thing that no longer works on current models is prefilling: ending messages with a partial assistant message to force how the reply starts. Opus 5.5, Sonnet 5.5, and the other recent models return a 400 error for it. Use a system prompt instruction or structured output instead.
Streaming Responses
For anything user-facing, stream the response so text appears as it's generated instead of after a long wait. The messages.stream() helper is the easiest way:
# stream_reply.py
import anthropic
client = anthropic.Anthropic()
with client.messages.stream(
model="claude-opus-5-5",
max_tokens=64000,
messages=[{"role": "user", "content": "Write a short README intro for a CSV-cleaning CLI."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message()
print(f"\n\nstop_reason={final.stop_reason}, output_tokens={final.usage.output_tokens}")
text_stream yields chunks of text as they arrive. get_final_message() returns the complete Message afterward, with the same content, stop_reason, and usage you'd get from create().
Streaming isn't only for UIs. Long generations over a non-streaming connection risk hitting HTTP timeouts, and the SDK refuses non-streaming requests that it estimates could run longer than about ten minutes, raising a ValueError that tells you to stream. Streaming and then calling get_final_message() gives you the timeout protection without handling individual events. That's also why the example uses a generous max_tokens=64000: with streaming, there's no reason to starve the response.
If you need finer control, iterate the stream itself to get typed events (content_block_start, content_block_delta, message_delta, and so on), or pass stream=True to messages.create() for a raw event iterator with no accumulation.
Effort and Thinking
Current Claude models can reason before answering. On Claude Opus 5.5, this thinking is always on and adapts to the difficulty of the request. You control how much work the model puts in with the effort setting:
# thinking.py
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={"type": "adaptive", "display": "summarized"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "Find the bug in this function: ..."}],
)
for block in response.content:
if block.type == "thinking":
print("[thinking]", block.thinking)
elif block.type == "text":
print(block.text)
output_config={"effort": ...}accepts"low","medium","high","xhigh", or"max". Higher effort means more thorough reasoning and more output tokens. On Opus 5.5 the default is"medium", so set it explicitly:"high"or"xhigh"for hard coding and agentic tasks,"low"for simple chat or classification.thinking={"type": "adaptive"}is the only thinking mode Opus 5.5 accepts, and it's what you get if you omit the parameter. Trying to disable thinking or set a fixedbudget_tokensreturns a 400 error. If you've seenbudget_tokensin older tutorials, that pattern is retired on current models; effort replaces it.displaycontrols whether you see a summary of the reasoning. By default, thinking blocks come back with empty text. Set"summarized"if you want to show or log a readable summary. You're billed for the thinking either way.
Thinking blocks are one more reason to append the full response.content to your history: they need to be sent back unchanged when the conversation continues.
Structured Output with Pydantic
When your code needs data rather than prose, don't parse free text with regular expressions. The SDK's messages.parse() takes a Pydantic model, constrains the response to its JSON schema, and returns a validated instance:
# extract_invoice.py
import anthropic
from pydantic import BaseModel
class Invoice(BaseModel):
vendor: str
total: float
currency: str
line_items: list[str]
client = anthropic.Anthropic()
email_body = """
Hi, attached is September's bill from Acme Cloud.
Compute: $199.00, Storage: $50.00. Total due: $249.00 USD.
"""
response = client.messages.parse(
model="claude-opus-5-5",
max_tokens=16000,
messages=[{"role": "user", "content": f"Extract the invoice details:\n{email_body}"}],
output_format=Invoice,
)
invoice = response.parsed_output
print(invoice.vendor, invoice.total, invoice.line_items)
response.parsed_output is an Invoice object with real types, so invoice.total + 10 is arithmetic, not string concatenation. If you're new to Pydantic models, they're the same BaseModel classes you'd use for request validation in FastAPI.
If you'd rather pass a raw JSON Schema, use messages.create() with output_config={"format": {"type": "json_schema", "schema": {...}}} and json.loads() the text block. Either way, the output is guaranteed to match the schema, which is far more reliable than asking for JSON in the prompt and hoping.
Tool Use: Letting Claude Call Your Functions
Tool use (also called function calling) lets the model request that your code run a function, such as looking up an order, querying a database, or calling a weather API, and then use the result in its answer. You describe each tool with a name, a description, and a JSON Schema for its inputs. The model never runs anything itself; it returns a tool_use block, your code executes the function, and you send back a tool_result.
The Manual Loop
# weather_tool.py
import json
import anthropic
client = anthropic.Anthropic()
tools = [
{
"name": "get_weather",
"description": "Get the current weather for a city. Returns temperature in Celsius.",
"strict": True,
"input_schema": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. Lisbon"},
},
"required": ["city"],
"additionalProperties": False,
},
}
]
def get_weather(city: str) -> dict:
# Replace with a real API call
return {"city": city, "temp_c": 21, "conditions": "sunny"}
messages = [{"role": "user", "content": "Should I take a jacket in Lisbon today?"}]
while True:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
tools=tools,
messages=messages,
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
break
tool_results = []
for block in response.content:
if block.type == "tool_use" and block.name == "get_weather":
result = get_weather(**block.input)
tool_results.append(
{
"type": "tool_result",
"tool_use_id": block.id,
"content": json.dumps(result),
}
)
messages.append({"role": "user", "content": tool_results})
print("".join(b.text for b in response.content if b.type == "text"))
The loop is the core of every agent:
- Send the conversation and the tool definitions.
- If
stop_reasonis"tool_use", find eachtool_useblock, run the matching function withblock.input(already parsed into a dict), and collect atool_resultfor each one, linked bytool_use_id. - Send all the results back in a single user message, then repeat until the model stops asking for tools.
A few practical notes:
"strict": Trueguarantees the tool input matches your schema exactly. Strict schemas need"additionalProperties": Falseand arequiredlist.- Descriptions matter. The model decides when and how to call a tool based on its name and description, so write them like documentation for a new colleague.
- Report failures, don't drop them. If a tool raises, return a
tool_resultwith"is_error": Trueand a short message so the model can recover or explain. - Validate before acting. Tool inputs come from the model, so treat them like user input. Check paths, IDs, and amounts before doing anything destructive.
- You can't force a specific tool on the newest models.
tool_choicevalues of"any"or a named tool return a 400 on Opus 5.5 and Sonnet 5.5. Leavetool_choiceat its default ("auto") and say in the prompt which tool to use.
The Tool Runner
Writing the loop yourself is worth doing once. For everyday use, the SDK includes a tool runner (currently in beta) that builds the schema from a typed function and its docstring, runs the loop, and calls your functions for you:
# tool_runner.py
import json
import anthropic
from anthropic import beta_tool
client = anthropic.Anthropic()
@beta_tool
def get_weather(city: str) -> str:
"""Get the current weather for a city.
Args:
city: City name, e.g. Lisbon.
"""
return json.dumps({"city": city, "temp_c": 21, "conditions": "sunny"})
runner = client.beta.messages.tool_runner(
model="claude-opus-5-5",
max_tokens=16000,
tools=[get_weather],
messages=[{"role": "user", "content": "Should I take a jacket in Lisbon today?"}],
)
final = runner.until_done()
print("".join(b.text for b in final.content if b.type == "text"))
@beta_tool reads the type hints and the Args: section of the docstring to generate the input schema, so good type hints do double duty here. until_done() runs the loop to completion and returns the final message. You can also iterate the runner (for message in runner:) to inspect or log each turn. Use @beta_async_tool with async def functions for async code.
Prompt Caching
Many applications send the same large prefix with every request: a long system prompt, a product manual, a codebase summary, tool definitions. Prompt caching stores that prefix on Anthropic's side so repeated requests read it from cache at a fraction of the normal input price (roughly a tenth), and with lower latency.
# cached_docs.py
from pathlib import Path
import anthropic
client = anthropic.Anthropic()
handbook = Path("handbook.md").read_text(encoding="utf-8")
def ask_handbook(question: str) -> str:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system=[
{"type": "text", "text": "Answer questions using only the handbook below."},
{"type": "text", "text": handbook, "cache_control": {"type": "ephemeral"}},
],
messages=[{"role": "user", "content": question}],
)
usage = response.usage
print(
f"cache write={usage.cache_creation_input_tokens} "
f"cache read={usage.cache_read_input_tokens} uncached={usage.input_tokens}"
)
return "".join(b.text for b in response.content if b.type == "text")
print(ask_handbook("How many vacation days do new hires get?"))
print(ask_handbook("What's the policy on remote work?"))
The cache_control marker on the handbook block says "cache everything up to and including this block". The first call writes the cache (cache_creation_input_tokens, billed at a small premium), and later calls within the cache lifetime read it (cache_read_input_tokens). The default lifetime is five minutes, refreshed each time it's used; {"type": "ephemeral", "ttl": "1h"} extends it to an hour. For the simplest setup, you can instead pass a top-level cache_control={"type": "ephemeral"} to messages.create(), which caches the last cacheable block automatically.
The rules that make or break caching:
- It's a prefix match. The request is rendered in order (tools, then system, then messages), and any change before the marker invalidates the cache. Put stable content first and anything that varies (the user's question, timestamps, IDs) after it.
- Watch for silent invalidators. A
datetime.now()in the system prompt, a dict serialized with unstable key order, or a tool list built in a different order each time will quietly prevent cache hits. - There's a minimum size. Prefixes shorter than a model-specific minimum (512 tokens on Opus 5.5) aren't cached, with no error.
- Verify it. If
cache_read_input_tokensstays at zero across repeated identical requests, something in your prefix is changing.
Async Requests
For web apps built on asyncio, or for running many independent requests concurrently, use AsyncAnthropic. It has the same methods, awaited:
# async_batch.py
import asyncio
import anthropic
client = anthropic.AsyncAnthropic()
semaphore = asyncio.Semaphore(5)
async def classify(review: str) -> str:
async with semaphore:
response = await client.messages.create(
model="claude-haiku-4-5",
max_tokens=256,
system="Classify the review as positive, negative, or mixed. Reply with one word.",
messages=[{"role": "user", "content": review}],
)
return "".join(b.text for b in response.content if b.type == "text").strip().lower()
async def main() -> None:
reviews = ["Loved it, fast shipping.", "Broke after a week.", "Good value, ugly color."]
labels = await asyncio.gather(*(classify(r) for r in reviews))
for review, label in zip(reviews, labels):
print(f"{label:<8} {review}")
asyncio.run(main())
The semaphore caps concurrency at five in-flight requests so a large batch doesn't immediately run into rate limits. This is also a good place for Haiku: short, simple classification where speed and cost matter more than depth. Streaming works the same way with async with client.messages.stream(...) and async for text in stream.text_stream.
If you have thousands of requests and don't need answers immediately, the Message Batches API (client.messages.batches.create(...)) processes them asynchronously at half the normal price, with results within 24 hours (often much sooner).
Error Handling, Retries, and Timeouts
The SDK raises typed exceptions, all subclasses of anthropic.APIError. Catch them from most specific to least specific so you can treat retryable and permanent failures differently:
# errors.py
import anthropic
client = anthropic.Anthropic(max_retries=3, timeout=60.0)
def summarize(text: str) -> str | None:
try:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
messages=[{"role": "user", "content": f"Summarize:\n{text}"}],
)
except anthropic.BadRequestError as exc:
print(f"Fix the request: {exc.message}")
return None
except anthropic.AuthenticationError:
print("Check ANTHROPIC_API_KEY")
return None
except anthropic.NotFoundError as exc:
print(f"Unknown model or endpoint: {exc.message}")
return None
except anthropic.RateLimitError as exc:
print(f"Rate limited; retry after {exc.response.headers.get('retry-after')}s")
return None
except anthropic.APIStatusError as exc:
print(f"API error {exc.status_code} (request {exc.request_id})")
return None
except anthropic.APIConnectionError:
print("Network problem reaching the API")
return None
if response.stop_reason == "refusal":
print("The model declined this request")
return None
if response.stop_reason == "max_tokens":
print("Warning: output was cut off at max_tokens")
return "".join(b.text for b in response.content if b.type == "text")
How the pieces fit:
| Exception | HTTP status | Retry? |
|---|---|---|
BadRequestError | 400 | No: the request itself is invalid |
AuthenticationError | 401 | No: bad or missing key |
PermissionDeniedError | 403 | No |
NotFoundError | 404 | No: wrong model ID or endpoint |
RateLimitError | 429 | Yes, after retry-after |
InternalServerError and other 5xx | 500, 529 | Yes, with backoff |
APIConnectionError / APITimeoutError | none | Yes |
The SDK already retries for you. By default it retries connection errors, 408, 409, 429, and 5xx responses twice with exponential backoff. Set max_retries on the client (or 0 to disable) rather than wrapping every call in your own retry loop. The default request timeout is 10 minutes; set timeout= in seconds on the client, or override both per call with client.with_options(timeout=20.0, max_retries=5).messages.create(...). Keep in mind that a timed-out request is retried too, so the worst-case wall-clock time is roughly timeout multiplied by the number of attempts.
Not every problem is an exception. A response with stop_reason == "max_tokens" succeeded but is incomplete, and a "refusal" means the model declined (for example, when a safety classifier flags the request), returned as a normal HTTP 200. Check stop_reason before trusting content. For apps where legitimate requests occasionally trip a classifier, the API offers server-side fallbacks that re-run a declined request on another model automatically; it's a beta feature enabled with client.beta.messages.create(..., betas=["server-side-fallback-2026-07-01"], fallbacks="default").
Log request IDs. Every response has response._request_id, and API errors expose exc.request_id. Include it in your logs; it's what Anthropic support needs to investigate a problem. For more on structured logging, see logging in Python.
Counting Tokens and Estimating Cost
To estimate cost before sending a request, or to check that a document fits in the context window, count tokens with the API rather than a third-party tokenizer:
import anthropic
client = anthropic.Anthropic()
count = client.messages.count_tokens(
model="claude-opus-5-5",
messages=[{"role": "user", "content": "Explain Python's GIL in two sentences."}],
)
print(count.input_tokens)
Pass the same system, tools, and messages you plan to send for an accurate number. Multiply by the model's per-token price for a cost estimate, and compare actual spend against response.usage after the fact.
A Production Checklist
- Keep the API key in the environment or a secrets manager, never in code.
- Default to streaming for anything long or user-facing; use
get_final_message()when you don't need the individual events. - Set
max_tokensgenerously and check forstop_reason == "max_tokens". - Set
effortexplicitly per route instead of relying on the default. - Use
messages.parse()or a JSON schema for machine-readable output. - Put stable content first and add
cache_controlfor any large repeated prefix, then confirm cache reads inusage. - Configure
max_retriesandtimeouton the client; catch typed exceptions most-specific first. - Log
usageand request IDs for every call. - Treat model output, especially tool inputs, as untrusted until validated.
Conclusion
The anthropic SDK keeps the Claude API compact: messages.create() for requests, messages.stream() for incremental output, messages.parse() for typed data, and a small loop (or the tool runner) for tool use. Around that core, a few habits make an integration production-ready: append full content blocks to conversation history, control depth with effort rather than retired thinking budgets, cache stable prefixes, and let the SDK's built-in retries handle transient failures while you handle the permanent ones.
Start with a single create() call, then add streaming, structured output, and tools as your app needs them. The patterns stay the same as you scale from a script to an agent.


