Short Answer
When a request triggers the provider’s safety policy, Claude does not return an error. The API still returns HTTP 200, but:contentis an empty array[]andoutput_tokensis0;stop_reasonisrefusal;stop_detailsnames the refusal category (for examplecyber) and carries a short English explanation.
message.content[0] directly will raise an IndexError. With the OpenAI-compatible format you get an empty string and finish_reason is refusal.
A refusal usually comes back within 1–2 seconds. Whether it is billed depends on the category: a refusal before any output in the cyber category (and some others) is not billed — see “Billing Rules” below.
What a Refusal Looks Like
Below is the same request that triggers a cyber refusal, called in four ways (measured on 2026-09-29, ids masked):- Native · Non-streaming
- Native · Streaming
- OpenAI-compatible · Non-streaming
- OpenAI-compatible · Streaming
POST /v1/messages, stream: false:stop_reason: "refusal". That partial output is incomplete and should be discarded.Refusal Categories
stop_details.category currently has five values:
category and explanation are null — that is a normal value. The explanation text may change at any time: display it, don’t string-match on it.
Common Triggers
category: "cyber" is the one developers hit most often. Claude has a real-time safeguard for cybersecurity requests, and all of these tasks can trigger it:
- Asking the model to find bugs in code, decide whether a piece of code “has a vulnerability”, or name the vulnerability type (CWE)
- Writing or completing exploit code, or penetration-testing steps
- Analyzing or rewriting malicious code
Billing Rules
Per the provider’s rules (as of September 2026; the provider may adjust them as it measures false-positive rates):usage still shows token counts — that is a count, not necessarily a charge.
How to Detect and Handle It
Check stop_reason before reading content
stop_reason == "refusal"; in the OpenAI-compatible format check finish_reason == "refusal". Read content only after ruling out a refusal.Record refusals as a result type
stop_details.category, instead of counting it as a failure to retry.Don't retry the same content
Reset context in multi-turn conversations
Review which content gets refused
category to see which tasks trigger them, then decide whether to send that content to a different model.- Anthropic SDK
- OpenAI SDK
What About Legitimate Security Research?
The Cyber Verification Program mentioned in the refusal text is the provider’s free application program for legitimate security work: after identity verification, “high-risk dual-use” tasks such as exploitation or offensive tooling development can be relaxed. “Prohibited uses” such as ransomware development or mass data exfiltration are blocked in all cases. The program is applied for by an organization admin on a first-party provider account. For third-party platforms the provider notes that “not all platforms participate”, and APIYI does not currently offer access to the program. So when calling through APIYI, for refused samples:- record them honestly as “refused” in your evaluation results, grouped by category;
- or process that content with a different model.
How It Differs from an OpenAI Refusal
FAQ
Am I billed for a refusal?
Am I billed for a refusal?
cyber, general_harms or null are not billed; bio, frontier_llm and reasoning_extraction bill the input; a mid-stream refusal bills the input plus the output already streamed. See “Billing Rules” above.Can refusals be turned off?
Can refusals be turned off?
Why are some samples in the same dataset refused and others not?
Why are some samples in the same dataset refused and others not?
Why is output_tokens 0 in usage on a refusal?
Why is output_tokens 0 in usage on a refusal?
max_tokens cutoff: the latter has stop_reason: "max_tokens" and output tokens equal to the limit you set.