A policy works best when its question describes one observable behavior. Write it so another person could decide whether a request or response matches without guessing your intent.
Name the behavior
Use a specific yes-or-no question rather than a broad category.
Set boundaries and exceptions
Say when the behavior counts and which similar cases should remain allowed:
Does the response expose a reusable authentication credential, excluding redacted values, placeholders, and clearly fictional examples?
If one question needs many unrelated exceptions, split it into separate policies. This makes each match easier to investigate.
Choose where to evaluate
- Use Requests when the decision must happen before the model runs.
- Use Responses when the risk lies in generated text or tool calls. In Block mode, output is held for evaluation, which increases response latency.
- Use both only when the same question makes sense for input and output. Otherwise, create two policies with questions tailored to each direction.
Test before blocking
- Publish in Monitor mode and test representative prompts in Playground.
- Review Decisions for expected matches, false positives, and failed checks.
- Adjust the question or sensitivity using those examples. Lower sensitivity values flag more borderline content.
- Switch to Block only when the observed results match the behavior you want.
Include ordinary allowed examples, borderline examples, and clear violations. Re-test after editing a policy; a saved draft does not affect traffic until it is published. Last modified on September 30, 2026