Function Registry

The function registry powers the eval policies and protocols and the safety / usefulness calculations consumed by ct game-metrics graph pareto. It handles argument declaration, CLI integration, and external module:function lookup.

For how policies and protocols fit into an eval, see Protocols and Policies. For the calculation-side registries used by Pareto plots, see Pareto Frontier.

Two Argument Shapes

Every registry builds a FunctionInfo (defined in function_info.py) for each registered function. There are two shapes for declaring arguments, each backed by a different function_info_factory:

ShapeFactoryArgs declared as…Used by
Args | Nonefunction_to_function_info (default — see function_info.py)a single BaseModelPython classsafety_calculations, usefulness_calculations
Direct kwargssignature_to_function_info (see signature_info.py)regular keyword parameterspolicies, protocols

Mixing shapes inside a registry is not supported. The factory is wired up in each registry's registry.py (function_info_factory=signature_to_function_info for the eval registries, omitted for the calculation registries).

Defining a Policy or Protocol (direct kwargs)

The eval registries, policies and protocols in protocols/eval/registry.py, declare configurable parameters directly as keyword arguments. The factory synthesizes a BaseModelPython args class from the function signature: each parameter's annotation and default become a field's type and default, and a pydantic Field(description=...) default documents it. Policies live in src/control_tower/protocols/eval/policies/ and protocols in src/control_tower/protocols/eval/protocols/.

from pydantic import Field from control_tower.protocols.eval.blue import Protocol from control_tower.protocols.eval.registry import protocols @protocols.register(name="trusted-monitoring") def trusted_monitoring( monitors: Monitors = Field(description=MONITORS_ARG_DESCRIPTION), sus_threshold: float | None = Field(default=None, description="Act above this score."), blocks: int = Field(default=0, ge=0), ) -> Protocol: """Ask the monitors about each action; above `sus_threshold`, block, then audit and stop.""" return TrustedMonitoring(monitors=monitors, sus_threshold=sus_threshold, blocks=blocks)

A policy is registered the same way with @policies.register(...) and returns a Policy; it also declares attempts_main_task, attempts_side_task and attempts_sneaky, which trajectory conversion records on every trajectory it produced. Registration is automatic via the decorator: the registry imports every non-underscore module under its discovery package on first lookup.

Defining a Safety or Usefulness Calculation (Args | None)

Safety / usefulness calculations are pluggable impls consumed by ct game-metrics graph pareto (and compute_pareto). Decorate with @safety_calculations.register() or @usefulness_calculations.register() and drop the file under metrics/safety_calculations/impls/ or metrics/usefulness_calculations/impls/. Each impl returns a single scalar plus metadata; curve-style sweeps and optional bootstrap CIs are handled by callers such as Pareto, not by the impls themselves.

from control_tower.utils.function_registry.base_model import BaseModelPython from control_tower.metrics.safety_calculations.registry import ( safety_calculations, ) from control_tower.metrics.safety_calculations.types import ( SafetyResult, ) from inspect_ai.log import EvalSample class TMSafetyArgs(BaseModelPython): mode: str = "optimized" audit_budget: float | None = None @safety_calculations.register() def tm_safety( honest: "list[EvalSample]", attack: "list[EvalSample]", args: TMSafetyArgs | None = None, ) -> "SafetyResult": if args is None: args = TMSafetyArgs() # ... fit rule and compute safety ... return SafetyResult(safety=..., metadata={"mode": args.mode})

Args classes must extend BaseModelPython (not plain Pydantic BaseModel). Reference the impl from YAML by its hyphenated name (e.g. tm-safety, defer-to-trusted).

Naming Convention

By default, the registry strips the _<identifier> suffix from the function name and converts underscores to hyphens. Each registry passes its own identifier (see registry.py):

RegistryidentifierAuto-strips suffixExample
policieseval-policynoneregister(name="honest") → honest
protocolseval-protocolnonenull_blue_team → null-blue-team
safety_calculationssafety_calculation_safety_calculationtm_safety → tm-safety
usefulness_calculationsusefulness_calculation_usefulness_calculationmain_task_success → main-task-success

No Python function name ends in _eval-policy, so an eval policy or protocol is named after its function unless you pass name="...": honest_policy would otherwise be honest-policy.

The docstring's first sentence becomes the description shown in help output.

Passing Arguments from the CLI

# Untrusted policies use -up (short for --untrusted-policy-arg) ct run eval --untrusted-policy sleep -up min_s=3 -up max_s=4 --task-file simple # Protocols use -bp (short for --protocol-arg) ct run eval \ --untrusted-policy simple-attack \ --protocol trusted-monitoring \ -bp 'monitors=[simple]' \ -bp sus_threshold=7.0 \ --task-file simple

Hyphens in keys are converted to underscores automatically. Unknown keys raise a validation error. Models are not policy arguments: an eval binds them to roles with --model-role (see Model Configuration). ct run sabotage-eval selects the same policies via --attack-policy / --honest-policy and a protocol via --protocol / --protocol-arg — see docs/sabotage-eval.md.

Discovering Available Arguments

ct protocols policies # List all policies ct protocols policies -v # List with all arguments ct protocols policies attack # Show args for a specific policy ct protocols protocols # List all protocols ct protocols models # List model aliases

Cost Tracking

CONTROL_TOWER_PRICING selects the rate source:

  • file (default) loads a generated JSON table (~/.cache/control-tower/model_rates.json, or CONTROL_TOWER_MODEL_RATES).
  • api loads the deployment's pricing collection over GET /api/pricing and needs a viewer token (ct login or CONTROL_TOWER_API_TOKEN).
  • off loads nothing. Model calls still run and bill, but every one is recorded as unknown cost, so use it only for runs that make no priced calls. The sabotage-eval preflight skips its pricing check under off.

With file or api, the table loads at Task construction, and a missing or empty table fails the eval before any sample starts. A model the loaded table has no row for stamps unknown cost, not $0.

uv run python scripts/generate_model_rates.py writes the file: vendors' own prices from LiteLLM's price map under bare names, and OpenRouter's catalog under each model's full openrouter/... name, which a lookup tries before the bare name. An OpenRouter row is the default route's price, so a run pinned to an endpoint that prices differently needs a manual row under the same key. A re-run keeps rows marked "manual": true and models the source has stopped publishing.

External Functions

Any CLI flag that accepts a policy or protocol name also accepts a module:function path, following the same convention as uvicorn app:factory. This lets you use policies and protocols defined in your own project without modifying control-tower.

ct run eval --untrusted-policy my_project.policies:my_attack --task-file simple

The function must have the same shape as a registered one (keyword arguments, returning a Policy or Protocol). External environments also work via --env my_project.envs:my_env_config, and --extra-src puts a directory of experiment code on the path without installing it.

Quick Reference

I want to...Command
List policiesct protocols policies
List protocolsct protocols protocols
See a policy's argsct protocols policies <name>
Override a model--model-role untrusted=opus / --model-role trusted=haiku
Use external function--untrusted-policy my.module:my_func