Skip to main content

Short Answer

When a request triggers the provider’s safety policy, Claude does not return an error. The API still returns HTTP 200, but:
  • content is an empty array [] and output_tokens is 0;
  • stop_reason is refusal;
  • stop_details names the refusal category (for example cyber) and carries a short English explanation.
This is model behavior, not an API failure. Code that reads message.content[0] directly will raise an IndexError. With the OpenAI-compatible format you get an empty string and finish_reason is refusal. A refusal usually comes back within 1–2 seconds. Whether it is billed depends on the category: a refusal before any output in the cyber category (and some others) is not billed — see “Billing Rules” below.

What a Refusal Looks Like

Below is the same request that triggers a cyber refusal, called in four ways (measured on 2026-09-29, ids masked):
POST /v1/messages, stream: false:
A refusal can also happen mid-stream: part of the text is streamed first and the response then ends with stop_reason: "refusal". That partial output is incomplete and should be discarded.

Refusal Categories

stop_details.category currently has five values: When a refusal does not map to a named category, both category and explanation are null — that is a normal value. The explanation text may change at any time: display it, don’t string-match on it.

Common Triggers

category: "cyber" is the one developers hit most often. Claude has a real-time safeguard for cybersecurity requests, and all of these tasks can trigger it:
  • Asking the model to find bugs in code, decide whether a piece of code “has a vulnerability”, or name the vulnerability type (CWE)
  • Writing or completing exploit code, or penetration-testing steps
  • Analyzing or rewriting malicious code
Batch evaluation and dataset distillation are the most exposed. Running a whole vulnerability dataset through the model item by item often gets a sizeable share of samples refused. A script that assumes content[0] always exists will crash on the refused item, which looks like “the API works only some of the time”.

Billing Rules

Per the provider’s rules (as of September 2026; the provider may adjust them as it measures false-positive rates): Billed or not, a refused request still counts toward your rate limits. usage still shows token counts — that is a count, not necessarily a charge.

How to Detect and Handle It

1

Check stop_reason before reading content

In the native format check stop_reason == "refusal"; in the OpenAI-compatible format check finish_reason == "refusal". Read content only after ruling out a refusal.
2

Record refusals as a result type

A refusal is a successful call, not a network error. For evaluation work, record it separately as “refused” together with stop_details.category, instead of counting it as a failure to retry.
3

Don't retry the same content

Resending the same content usually gets the same refusal, still uses up rate limits, and for some categories is billed every time.
4

Reset context in multi-turn conversations

After a turn is refused, remove or rewrite that turn, or clear the history, before continuing. Without a reset, later requests will keep being refused.
5

Review which content gets refused

Group refusals by category to see which tasks trigger them, then decide whether to send that content to a different model.
If you need the refusal category, call the native /v1/messages format. The OpenAI-compatible format keeps only finish_reason: "refusal" and has no stop_details.

What About Legitimate Security Research?

The Cyber Verification Program mentioned in the refusal text is the provider’s free application program for legitimate security work: after identity verification, “high-risk dual-use” tasks such as exploitation or offensive tooling development can be relaxed. “Prohibited uses” such as ransomware development or mass data exfiltration are blocked in all cases. The program is applied for by an organization admin on a first-party provider account. For third-party platforms the provider notes that “not all platforms participate”, and APIYI does not currently offer access to the program. So when calling through APIYI, for refused samples:
  • record them honestly as “refused” in your evaluation results, grouped by category;
  • or process that content with a different model.
APIYI does not and cannot adjust the provider’s safety policy.

How It Differs from an OpenAI Refusal

For what an OpenAI refusal looks like, see What Does an OpenAI Model Refusal Look Like?.

FAQ

It depends on the category and timing. Pre-output refusals in cyber, general_harms or null are not billed; bio, frontier_llm and reasoning_extraction bill the input; a mid-stream refusal bills the input plus the output already streamed. See “Billing Rules” above.
No. Refusals are decided by the provider’s model under its safety policy; APIYI cannot turn them off or change how strict they are. Record refused content as a refusal result, or process it with a different model.
The safeguard judges each request on its own content. The code snippet itself and how the prompt is phrased both affect the outcome, so usually only part of a dataset is refused. In our tests, resending the same refused sample gave a consistent result.
The model stopped before generating anything, so output is 0 and only input is counted. This is also how you tell a refusal from a max_tokens cutoff: the latter has stop_reason: "max_tokens" and output tokens equal to the limit you set.

Claude Response Handling

Streaming and non-streaming response structure, stop_reason values

What Does an OpenAI Model Refusal Look Like?

What a GPT refusal looks like and how to detect it

How is content safety and compliance ensured?

Platform content safety and compliance policy