All documentation

Guardrails

Proxyma checks what people send, what the agent retrieves, and what it answers, and holds back anything that matches your rules. This page covers each check, what a person sees when one fires, and what guardrails do not stop.

Guardrails are on when you install Proxyma

All three checkpoints and the four pattern checks are enabled. They need no configuration, no AI provider and no internet connection. The layers that use a model are off until you turn them on and choose a model.

Where to find them

Settings → Guardrails, in both the desktop app and Proxyma Enterprise Server. Settings apply to the whole deployment, not per user. In Server Mode, only administrators can open the page. Administrators can also add and remove policy documents and clear Logs; only a super administrator can change the settings.

The three checkpoints

Every turn passes all three, and each has its own default.

CheckpointWhat it protectsDefault
What people send Your deployment and your AI provider. It runs before the message is stored or sent to the chat model, so a held-back message is never saved or answered. It leaves the machine only if you turn on the model layers below. Block the turn
What the agent reads The agent. Search results, web pages, connector documents, issue text and every tool result, including connectors added later, are checked before the model sees them. Remove and continue
What the agent answers The reader. Pattern checks run while the answer is written, so a match is never shown; the last few hundred characters of each answer appear when it finishes rather than word by word. Remove and continue

What "held back" means at each one

  • A message is refused. The turn does not happen and nothing is stored.
  • A retrieved source is withheld on its own. The agent is told it was held back and continues with everything else; this never ends the conversation.
  • An answer is rewritten where it is stored. Only its final reply is: text the agent wrote between tool calls earlier in the same turn stays in the conversation history.

What the checks look for

These are pattern checks: they recognize shapes, not meaning. They use no model and no network.

CheckWhat it matchesWhere it runs
Credentials API keys, access tokens, private keys, and long random-looking values assigned to a name like password or secret All three
Personal data Email addresses, phone numbers, payment card numbers Messages and answers
Prompt injection Text that tries to talk the agent out of its instructions Messages and retrieved content
Server file paths Full paths on the machine running Proxyma, appearing in an answer Answers

Personal data is not checked in retrieved content; it is still checked in the answer. Server paths are not checked in retrieved content either.

Text is normalized before matching, so a credential split with an invisible character, or a word spelled with lookalike letters from another alphabet, is still matched. Text genuinely written in another alphabet is left alone.

What somebody sees

A held-back message or answer is replaced by a short notice naming the categories that fired, such as secret, pii, prompt-injection, server-path or system-prompt-leak. A removal keeps the rest of the text and marks the gap, for example aws_key: [redacted: secret].

A withheld source appears as a failed step in the agent's trace, with the categories that fired.

The text that matched is never shown, logged or counted - unless you ask for it

Notices name categories, and logs name rules and categories. Neither carries the message, the matched span or the policy clause. If you turn on diagnostics, they count the checks, their verdicts and how long they took - never any text. There is no setting to change this.

The Logs section is the one opt-in exception. With Keep the matched text turned on (at the foot of the Guardrails page; off by default), a decision in that section shows the characters each rule matched. These spans are held in memory, capped at 200 decisions and 120 characters each, and cleared when Proxyma restarts. They are never written to the database, a log or diagnostics. Any administrator who can open the Settings screen can read them. On a server install the Logs section and its endpoint are restricted to administrators; an ordinary signed-in user cannot reach either.

This setting does not change what is detected, blocked or redacted.

Watch before you enforce

  1. Set Mode to Monitor only. Every check still runs, nothing is blocked or removed - not even while an answer is being written - and what would have happened is recorded.
  2. Open Logs, at the foot of the same page. It shows how many checks ran - one per message, per answer and per tool result - and how many were allowed, redacted and blocked; how many would have been withheld if enforcing; the time guardrails added per turn in milliseconds; which categories fired; and the last few decisions with their stage, verdict and timing.
  3. Open a decision to see the rules behind it, for example secret.aws-access-key rather than secret. A decision made by the model lists llm.category, or policy. followed by the document's id.
  4. When the results look right, set Mode to Enforce.

Logs are held in memory and cleared when Proxyma restarts.

Your own rules

Turn on Judge meaning with a model, then Rules from your documents under it.

Rules from your documents

Upload your policies - a handbook, an acceptable-use policy, a data-handling standard - as Markdown, PDF, Word, Excel or PowerPoint files, up to 20 MB each. Each message is matched by meaning against the relevant clauses, and those clauses are sent to the model with it. Turn on Also check answers to match answers against them too.

Policy documents live in their own index. No user question can retrieve one.

The model that judges

With Judge meaning with a model on, every message and answer the pattern checks have not already blocked gets a second model call, using its own Guardrail Provider and Guardrail Model, never your chat model. It judges against a built-in list - violence, self-harm, sexual content involving minors, instructions for weapons or explosives, harassment or hate speech - plus any clauses matched from your documents. Choose something small and fast. Nothing is judged until a model is chosen; the pattern checks keep working either way.

What leaves the machine when you turn this on

The message or answer being judged, and the policy clauses that matched it, are sent to the provider you chose for the guardrail model. With Rules from your documents on, the text being judged is also sent to your embedding provider to find those clauses. Point both at local models and nothing leaves the machine. The pattern checks send nothing anywhere.

By default, a message goes through when that model cannot be reached. Turn on Refuse when the model cannot be reached to reverse this. With it off, a provider outage removes this layer; with it on, a provider outage stops people working.

Two checks that watch for a pattern

Neither of these reads a message.

Detect a leaked system prompt

Proxyma hides an unguessable marker in the agent's instructions and refuses any answer that contains it. It needs no model and cannot mistake anything else for a leak. Off by default.

It catches an answer that quotes the instructions. It does not catch one that paraphrases them.

Slow down repeated attempts

When on, somebody refused several times in a short window is paused: their messages are refused without being checked, and they are told it is not about what they wrote. Attempts are counted per person where there are accounts and per conversation otherwise, and forgotten when Proxyma restarts.

Off by default. Leave it off unless you need it: somebody pasting error logs that contain tokens can trip a check several times in a row and be locked out.

What guardrails are not

  • They are not access control. A guardrail is a text filter. Who may read which source is decided by connector permissions and, in Server Mode, by accounts and roles.
  • The prompt-injection patterns are a speed bump, not a lock. They catch copied-and-pasted attempts, not a careful rewording. Proxyma's own rules tell the agent never to treat retrieved content as an instruction; these patterns sit underneath that.
  • The retrieval check lets a source through if the check itself fails.
  • The Logs section is not an audit log. It is in memory and starts empty after a restart.

Getting help

Ask at contact@proxyma.ai. If a guardrail is refusing work it should not, switch to monitor mode first to keep people working while you investigate.