Strategies

Base class for guardrail strategies.

Each GuardrailStrategy has exactly one concrete subclass, registered in smarter.apps.guardrail.services.strategies.registry. A strategy’s job is narrow: given a text segment and the Guardrail that owns it, decide whether it triggers, with what confidence, and, for the strategies that can, where. What happens next (block, redact, flag and so on) is the job of smarter.apps.guardrail.services.actions.

class smarter.apps.guardrail.services.strategies.base.BaseGuardrailStrategy[source]

Bases: ABC

Abstract base for all guardrail strategies.

Instances are stateless and safe to reuse across requests.

abstractmethod evaluate(*, segment, guardrail, context)[source]

Evaluate a single text segment against a single Guardrail.

Parameters:
  • segment (TextSegment) – The text segment.

  • guardrail (Guardrail) – The Guardrail, which supplies the strategy’s configuration.

  • context (StrategyContext) – Ambient evaluation context.

Return type:

StrategyMatch

Returns:

Whether the segment triggered the guardrail, and how.

Raises:
  • GuardrailConfigError – If the Guardrail’s configuration is invalid for the strategy. A misconfigured guardrail must fail loudly, rather than silently pass.

  • GuardrailProviderError – If a call to an LLM provider fails.

static threshold(guardrail)[source]

Return the Guardrail’s confidence threshold, or the default.

Return type:

float

class smarter.apps.guardrail.services.strategies.base.StrategyContext(**data)[source]

Bases: BaseModel

Ambient information a strategy may need beyond the text itself.

Variables:
  • stage – Which side of the model call is being evaluated.

  • request_uid – A caller-supplied identifier for correlating strategy calls with a request.

model_config: ClassVar[ConfigDict] = {'arbitrary_types_allowed': True}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

request_uid: str | None
stage: GuardrailStage
class smarter.apps.guardrail.services.strategies.base.StrategyMatch(**data)[source]

Bases: BaseModel

What a strategy hands back for a single segment.

Variables:
  • triggered – Whether the segment matched the guardrail.

  • confidence – A confidence score in [0.0, 1.0]. The deterministic strategies report 1.0 when they trigger.

  • matches – Where the segment matched, for the strategies that locate what they match.

  • rationale – A human-readable explanation, for events and logs.

confidence: float | None
matches: list[GuardrailMatch]
model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

rationale: str | None
triggered: bool

Maps each GuardrailStrategy to its strategy.

The strategies are stateless singletons. The scored strategies get their client from get_client() when they evaluate, so that each Guardrail may name its own provider. configure_clients(), re-exported here, replaces the client factory, e.g. with a fake in tests.

smarter.apps.guardrail.services.strategies.registry.configure_clients(client_factory=None)[source]

Replace the factory of the scored strategies’ clients, or restore the default.

Parameters:

client_factory (Optional[Callable[[Guardrail], GuardrailClient]]) – A function of a Guardrail that returns its client, or None to restore default_client_for().

Return type:

None

smarter.apps.guardrail.services.strategies.registry.get_client(guardrail)[source]

Return the client of the Guardrail’s LLM provider, from the configured factory.

Return type:

GuardrailClient

smarter.apps.guardrail.services.strategies.registry.get_strategy(strategy)[source]

Return the strategy for a GuardrailStrategy value.

Raises:

GuardrailStrategyNotImplementedError – If the strategy is unknown.

Return type:

BaseGuardrailStrategy

regex: guardrail.pattern is a single Python regular expression.

guardrail.config["flags"] may contain any of IGNORECASE, MULTILINE and DOTALL. It defaults to ["IGNORECASE"].

class smarter.apps.guardrail.services.strategies.regex_strategy.RegexStrategy[source]

Bases: BaseGuardrailStrategy

Match a text segment against a regular expression.

evaluate(*, segment, guardrail, context)[source]

Find every match of guardrail.pattern in the segment.

Return type:

StrategyMatch

smarter.apps.guardrail.services.strategies.regex_strategy.compile_pattern(pattern, flags)[source]

Compile and cache a regex pattern with the given named flags.

Raises:

GuardrailConfigError – If a flag is unknown, or the pattern does not compile.

Return type:

Pattern

smarter.apps.guardrail.services.strategies.regex_strategy.compiled_pattern(guardrail)[source]

Return the Guardrail’s compiled pattern.

Return type:

Pattern

keyword: guardrail.config["keywords"] is a list of words and phrases.

guardrail.config["caseSensitive"] defaults to false, and guardrail.config["wholeWord"] to true. Whitespace within a phrase matches any whitespace, so that a phrase split across lines still matches.

class smarter.apps.guardrail.services.strategies.keyword_strategy.KeywordStrategy[source]

Bases: BaseGuardrailStrategy

Match a text segment against a list of words and phrases.

evaluate(*, segment, guardrail, context)[source]

Find every occurrence of the Guardrail’s keywords in the segment.

Return type:

StrategyMatch

smarter.apps.guardrail.services.strategies.keyword_strategy.compile_keywords(keywords, case_sensitive, whole_word)[source]

Compile and cache a keyword list into a single alternation, longest keywords first.

Return type:

Pattern

Built-in detectors of personal data and secrets, for the detector strategy.

Each detector is a regular expression, and, where the data has a checksum, a validator, so that e.g. a 16 digit order number is not mistaken for a credit card number. The detectors favor precision over recall: they match well-formed data, and do not attempt to find every possible formatting of it.

Note

Experimental. The Guardrail was designed and coded by Claude Code (Anthropic’s Claude Opus 5.5), with Lawrence McDaniel as co-author. It is experimental, and will be documented.

smarter.apps.guardrail.services.strategies.detectors.DETECTORS: dict[str, Detector] = {'api_key': Detector(name='api_key', pattern=re.compile('(?<![\\w-])(?:sk-(?:proj-|ant-)?[A-Za-z0-9_-]{20,}|gh[pousr]_[A-Za-z0-9]{36,}|github_pat_[A-Za-z0-9_]{22,}|xox[abposr]-[A-Za-z0-9-]{10,}|AIza[0-9A-Za-z_-]{35}|(?:sk|rk)_(?:live|test)_[0-9A-Za-z]{24,}), validator=None), 'aws_access_key': Detector(name='aws_access_key', pattern=re.compile('(?<![A-Z0-9])(?:AKIA|ASIA|ABIA|ACCA)[A-Z0-9]{16}(?![A-Z0-9])'), validator=None), 'credit_card': Detector(name='credit_card', pattern=re.compile('(?<![\\d-])\\d(?:[ -]?\\d){12,18}(?![\\d-])'), validator=<function credit_card_valid>), 'email': Detector(name='email', pattern=re.compile('(?<![\\w.+-])[A-Za-z0-9._%+-]+@[A-Za-z0-9-]+(?:\\.[A-Za-z0-9-]+)*\\.[A-Za-z]{2,}(?![\\w-])'), validator=None), 'iban': Detector(name='iban', pattern=re.compile('(?<![A-Za-z0-9])[A-Z]{2}\\d{2}(?: ?[A-Z0-9]){11,30}(?![A-Za-z0-9])'), validator=<function iban_valid>), 'ip_address': Detector(name='ip_address', pattern=re.compile('(?<![\\w.:])(?:(?:\\d{1,3}\\.){3}\\d{1,3}|(?:[0-9A-Fa-f]{1,4}:){2,7}[0-9A-Fa-f]{0,4}(?::[0-9A-Fa-f]{1,4})*)(?![\\w.:])'), validator=<function ip_address_valid>), 'jwt': Detector(name='jwt', pattern=re.compile('(?<![\\w-])eyJ[A-Za-z0-9_-]{8,}\\.eyJ[A-Za-z0-9_-]{8,}\\.[A-Za-z0-9_-]{8,}(?![\\w-])'), validator=None), 'password_assignment': Detector(name='password_assignment', pattern=re.compile('(?i)\\b(?:password|passwd|pwd|passphrase|secret|client_secret|api[_-]?key|access[_-]?token|auth[_-]?token)[ \\t]*[:=][ \\t]*[\\"\']?(?P<value>[^\\s\\"\',;]{6,})', re.IGNORECASE), validator=None), 'phone_number': Detector(name='phone_number', pattern=re.compile('(?<![\\w+])(?:\\+\\d{1,3}[ .-]?(?:\\(\\d{1,4}\\)|\\d{1,4})(?:[ .-]?\\d{2,4}){2,4}|(?:1[ .-])?(?:\\(\\d{3}\\)\\s?|\\d{3}[ .-])\\d{3}[ .-]\\d{4})(?![\\w])'), validator=None), 'private_key': Detector(name='private_key', pattern=re.compile('-----BEGIN (?:[A-Z]+ )*PRIVATE KEY(?: BLOCK)?-----[\\s\\S]*?(?:-----END (?:[A-Z]+ )*PRIVATE KEY(?: BLOCK)?-----|\\Z)'), validator=None), 'us_ssn': Detector(name='us_ssn', pattern=re.compile('(?<![\\d-])(?!000|666|9\\d\\d)\\d{3}[- ](?!00)\\d{2}[- ](?!0000)\\d{4}(?![\\d-])'), validator=None)}

The built-in detectors, by name.

{
    'api_key': Detector(name='api_key', pattern=re.compile('(?<![\\w-])(?:sk-(?:proj-|ant-)?[A-Za-z0-9_-]{20,}|gh[pousr]_[A-Za-z0-9]{36,}|github_pat_[A-Za-z0-9_]{22,}|xox[abposr]-[A-Za-z0-9-]{10,}|AIza[0-9A-Za-z_-]{35}|(?:sk|rk)_(?:live|test)_[0-9A-Za-z]{24,}), validator=None),
    'aws_access_key': Detector(name='aws_access_key', pattern=re.compile('(?<![A-Z0-9])(?:AKIA|ASIA|ABIA|ACCA)[A-Z0-9]{16}(?![A-Z0-9])'), validator=None),
    'credit_card': Detector(name='credit_card', pattern=re.compile('(?<![\\d-])\\d(?:[ -]?\\d){12,18}(?![\\d-])'), validator=<function credit_card_valid at 0x11e04c900>),
    'email': Detector(name='email', pattern=re.compile('(?<![\\w.+-])[A-Za-z0-9._%+-]+@[A-Za-z0-9-]+(?:\\.[A-Za-z0-9-]+)*\\.[A-Za-z]{2,}(?![\\w-])'), validator=None),
    'iban': Detector(name='iban', pattern=re.compile('(?<![A-Za-z0-9])[A-Z]{2}\\d{2}(?: ?[A-Z0-9]){11,30}(?![A-Za-z0-9])'), validator=<function iban_valid at 0x11e04ca40>),
    'ip_address': Detector(name='ip_address', pattern=re.compile('(?<![\\w.:])(?:(?:\\d{1,3}\\.){3}\\d{1,3}|(?:[0-9A-Fa-f]{1,4}:){2,7}[0-9A-Fa-f]{0,4}(?::[0-9A-Fa-f]{1,4})*)(?![\\w.:])'), validator=<function ip_address_valid at 0x11e04cb80>),
    'jwt': Detector(name='jwt', pattern=re.compile('(?<![\\w-])eyJ[A-Za-z0-9_-]{8,}\\.eyJ[A-Za-z0-9_-]{8,}\\.[A-Za-z0-9_-]{8,}(?![\\w-])'), validator=None),
    'password_assignment': Detector(name='password_assignment', pattern=re.compile('(?i)\\b(?:password|passwd|pwd|passphrase|secret|client_secret|api[_-]?key|access[_-]?token|auth[_-]?token)[ \\t]*[:=][ \\t]*[\\"\']?(?P<value>[^\\s\\"\',;]{6,})', re.IGNORECASE), validator=None),
    'phone_number': Detector(name='phone_number', pattern=re.compile('(?<![\\w+])(?:\\+\\d{1,3}[ .-]?(?:\\(\\d{1,4}\\)|\\d{1,4})(?:[ .-]?\\d{2,4}){2,4}|(?:1[ .-])?(?:\\(\\d{3}\\)\\s?|\\d{3}[ .-])\\d{3}[ .-]\\d{4})(?![\\w])'), validator=None),
    'private_key': Detector(name='private_key', pattern=re.compile('-----BEGIN (?:[A-Z]+ )*PRIVATE KEY(?: BLOCK)?-----[\\s\\S]*?(?:-----END (?:[A-Z]+ )*PRIVATE KEY(?: BLOCK)?-----|\\Z)'), validator=None),
    'us_ssn': Detector(name='us_ssn', pattern=re.compile('(?<![\\d-])(?!000|666|9\\d\\d)\\d{3}[- ](?!00)\\d{2}[- ](?!0000)\\d{4}(?![\\d-])'), validator=None),
}
class smarter.apps.guardrail.services.strategies.detectors.Detector(name, pattern, validator=None)[source]

Bases: object

A built-in detector.

Variables:
  • name – The detector’s name, as in the manifest’s detectors.

  • pattern – The regular expression. If it has a group named value, only that group is matched, e.g. the password of password=hunter22, rather than the assignment.

  • validator – An optional function that confirms a match.

__init__(name, pattern, validator=None)
name: str
pattern: Pattern
validator: Callable[[str], bool] | None = None
smarter.apps.guardrail.services.strategies.detectors.credit_card_valid(text)[source]

Return whether text is a plausible credit card number: 13 to 19 digits, not all the same, that pass Luhn.

Return type:

bool

smarter.apps.guardrail.services.strategies.detectors.detect(text, detectors)[source]

Return the matches of the named detectors in text.

Overlapping matches are merged, keeping the longest, so that e.g. a credit card number is not also reported as a phone number.

Parameters:
  • text (str) – The text to scan.

  • detectors (list[str]) – The names of the detectors to run.

Return type:

list[GuardrailMatch]

Returns:

The matches, in order of their position in the text.

Raises:

KeyError – If a detector name is unknown.

smarter.apps.guardrail.services.strategies.detectors.iban_valid(text)[source]

Return whether text passes the ISO 13616 mod-97 checksum of IBANs.

Return type:

bool

smarter.apps.guardrail.services.strategies.detectors.luhn_valid(digits)[source]

Return whether a string of digits passes the Luhn checksum of credit card numbers.

Return type:

bool

detector: guardrail.config["detectors"] names built-in detectors of personal data and secrets.

class smarter.apps.guardrail.services.strategies.detector_strategy.DetectorStrategy[source]

Bases: BaseGuardrailStrategy

Match a text segment with built-in, validated detectors.

See detectors.

evaluate(*, segment, guardrail, context)[source]

Find everything that the Guardrail’s detectors detect in the segment.

Return type:

StrategyMatch

semantic: triggers when the segment’s embedding is similar to any of.

guardrail.config["referenceTexts"].

The similarity is the cosine similarity of the embeddings, and the guardrail triggers when the best similarity reaches guardrail.threshold. The reference texts’ embeddings are cached, so that each prompt embeds only the segment.

class smarter.apps.guardrail.services.strategies.semantic_strategy.SemanticStrategy[source]

Bases: BaseGuardrailStrategy

Match a text segment by its embedding’s similarity to reference texts.

evaluate(*, segment, guardrail, context)[source]

Compare the segment’s embedding with each reference text’s embedding.

Return type:

StrategyMatch

smarter.apps.guardrail.services.strategies.semantic_strategy.cosine_similarity(a, b)[source]

Return the cosine similarity of two vectors, or 0.0 if either has no magnitude.

Raises:

GuardrailProviderError – If the vectors have different lengths.

Return type:

float

moderation: an LLM provider’s moderation model, e.g. OpenAI’s omni-moderation-latest.

Triggers when the score of any of guardrail.config["categories"], or of any category if it is empty, reaches guardrail.threshold, or when the moderation model flags one of them.

class smarter.apps.guardrail.services.strategies.moderation_strategy.ModerationStrategy[source]

Bases: BaseGuardrailStrategy

Match a text segment with a moderation model.

evaluate(*, segment, guardrail, context)[source]

Moderate the segment, and compare the scores of the Guardrail’s categories with its threshold.

Return type:

StrategyMatch

llm_judge: an LLM judges the segment, with guardrail.config["judgePrompt"].

The prompt’s {text} placeholder is replaced with the segment’s text. The LLM must reply with a JSON object: {"triggered": true or false, "confidence": 0.0 to 1.0, "rationale": "..."}. The guardrail triggers when the judge says so, with a confidence that reaches guardrail.threshold. The judge runs at guardrail.config["temperature"], 0 by default, so that its verdicts are as deterministic as possible.

class smarter.apps.guardrail.services.strategies.llm_judge_strategy.LLMJudgeStrategy[source]

Bases: BaseGuardrailStrategy

Match a text segment by asking an LLM to judge it.

evaluate(*, segment, guardrail, context)[source]

Render the judge prompt for the segment, and run it.

Return type:

StrategyMatch

The LLM provider calls of the scored strategies.

The semantic, moderation and llm_judge strategies each call an LLM provider: for embeddings, a moderation model, and a chat completion, respectively. They do so through a GuardrailClient, so that they are independent of any one provider, and testable with a fake.

OpenAIGuardrailClient implements it with the OpenAI Python SDK, for any OpenAI-compatible Smarter Provider. default_client_for() returns one for the Provider that a Guardrail names in its provider field, openai by default, as its owner may read it.

class smarter.apps.guardrail.services.strategies.clients.GuardrailClient(*args, **kwargs)[source]

Bases: Protocol

The LLM provider calls of the scored strategies.

__init__(*args, **kwargs)
embed(texts, *, model)[source]

Return an embedding vector for each text.

Return type:

list[list[float]]

judge(prompt, *, model, temperature=0.0)[source]

Run a judge prompt, and return its verdict.

Return type:

JudgeVerdict

moderate(text, *, model)[source]

Return a moderation model’s category scores for text.

Return type:

ModerationResult

class smarter.apps.guardrail.services.strategies.clients.JudgeVerdict(triggered, confidence, rationale='')[source]

Bases: object

The result of an LLM judge call.

Variables:
  • triggered – Whether the judge considered the guardrail’s condition met.

  • confidence – The judge’s confidence, from 0 to 1.

  • rationale – The judge’s explanation.

__init__(triggered, confidence, rationale='')
confidence: float
rationale: str = ''
triggered: bool
class smarter.apps.guardrail.services.strategies.clients.ModerationResult(scores=<factory>, flagged=<factory>)[source]

Bases: object

The result of a moderation call.

Variables:
  • scores – The score of each category, from 0 to 1, e.g. {"hate": 0.02}.

  • flagged – The categories that the moderation model flagged.

__init__(scores=<factory>, flagged=<factory>)
flagged: list[str]
scores: dict[str, float]
class smarter.apps.guardrail.services.strategies.clients.OpenAIGuardrailClient(api_key, base_url=None)[source]

Bases: object

A GuardrailClient for an OpenAI-compatible API.

Parameters:
  • api_key (str) – The API key.

  • base_url (Optional[str]) – The API’s base URL, or None for OpenAI’s.

__init__(api_key, base_url=None)[source]
embed(texts, *, model)[source]
Return type:

list[list[float]]

judge(prompt, *, model, temperature=0.0)[source]
Return type:

JudgeVerdict

moderate(text, *, model)[source]
Return type:

ModerationResult

smarter.apps.guardrail.services.strategies.clients.configure_clients(client_factory=None)[source]

Replace the factory of the scored strategies’ clients, or restore the default.

Parameters:

client_factory (Optional[Callable[[Guardrail], GuardrailClient]]) – A function of a Guardrail that returns its client, or None to restore default_client_for().

Return type:

None

smarter.apps.guardrail.services.strategies.clients.default_client_for(guardrail)[source]

Return a client for the Provider that the Guardrail names, as its owner may read it.

Clients are cached by the Provider’s id, its update time and its API key, so that a changed Provider gets a new client.

Raises:

GuardrailConfigError – If the Provider does not exist, or has no API key.

Return type:

GuardrailClient

smarter.apps.guardrail.services.strategies.clients.get_client(guardrail)[source]

Return the client of the Guardrail’s LLM provider, from the configured factory.

Return type:

GuardrailClient

smarter.apps.guardrail.services.strategies.clients.parse_verdict(content)[source]

Parse an LLM judge’s reply: a JSON object with triggered, confidence and rationale.

Raises:

GuardrailProviderError – If the reply is not such a JSON object.

Return type:

JudgeVerdict