ephor.ragbaz.cc
Work in Progress · Integration Roadmap

WeftMark+Ephor

WeftMark governs who touches your code and whether the evidence holds up. Ephor governs whether agent actions were compliant and provably reviewed. Together they close the loop between engineering provenance and regulatory accountability.


Two products, one oversight gap

What each one does

Both systems deal with multi-agent work. Neither alone is sufficient for teams building high-risk AI under the EU AI Act.

WeftMark · Control Plane

Engineering provenance

The multi-agent software coordination layer. WeftMark tracks Change Sets (the envelope around an agent's scope and branch), an evidence ledger with typed proof (unit tests, CI, security review), semantic scope locks that detect behavioral contract collisions, portable handoff records, and reviewer-facing readiness decisions. Vendor-neutral: Claude Code, Codex, OpenCode, and Ollama-backed workers are all just actors in the ledger.

Who was allowed to change this? What exactly changed? What evidence says it is correct? Who decided it is ready to merge?

Ephor · Governance Layer

Regulatory compliance

The AI governance magistrate. Every agent action passes through Ephor's policy gate before execution. High-risk actions are held pending human approval (Art 14 HITL) with SLA timers enforcing safe defaults on timeout. Every entry is appended to a SHA-256 hash chain — tamper-evident, verifiable, append-only. Evidence packs export the full compliance record for conformity assessment (Art 11, Art 47).

Was this agent action permitted by policy? Was a human in the loop? Can we produce an immutable audit trail that proves EU AI Act adherence?

Without WeftMark, Ephor audits individual actions but cannot explain why an agent had the authority to act — what scope it declared, what evidence it produced, or whether a reviewer signed off on the Change Set before the agent started operating.

Without Ephor, WeftMark tracks engineering evidence but cannot enforce compliance policy, hold actions pending human review, or produce the tamper-evident audit chain that Art 12 requires for high-risk AI systems. Review is a state in a ledger, not a legal record.

Together, the two systems produce something neither can deliver alone: a continuous, verifiable chain from declared agent scope → policy gate → HITL decision → evidence record → reviewer sign-off → immutable audit entry. Engineering provenance and regulatory governance as one observable pipeline.


Integration

Five surfaces where the systems connect

These are the natural join points — each one is buildable independently and useful on its own, but together they form a closed loop.

01

Change Set events → Ephor capture

When WeftMark opens a Change Set — a worker is assigned scope and begins operating on a branch — this is an auditable agent action in Ephor's terms. A WeftMark-Ephor adapter posts lifecycle events (changeset.open, scope.claimed, changeset.close) to Ephor's POST /capture endpoint. Agent class, session ID, declared scopes, and semantic locks all travel with the entry. The audit chain grows a record not just of what the agent did, but of what it was authorised to do before it started.

weftmark/adapters/ephor.py POST /capture
02

Semantic scopes as risk signals

WeftMark's semantic scopes declare behavioral contract boundaries — a Change Set touching contract:tenant-authentication or contract:payment-routing is not just an edit to a file, it's a change to a security or financial protocol. These are natural risk inputs to Ephor's policy engine. A scope-to-risk mapping config lets teams declare that any contract: scope claim elevates the Ephor risk level to high, automatically triggering Art 14 HITL before the Change Set can proceed. Scope risk is no longer a judgment call made after the fact — it's enforced at execution time.

risk:high on contract: scopes Art 14 · HITL
03

Evidence kind: ephor:governance

WeftMark already has a typed evidence model — unit tests, CI, security review, benchmarks. Adding an ephor:governance evidence kind lets a Change Set's evidence policy require Ephor compliance as a verifiable precondition of merge. A Change Set targeting a high-risk scope must carry an ephor:governance entry in PASSED state before it can advance to READY. The evidence record cites the Ephor chain hash, so WeftMark's ledger points to the tamper-evident compliance record as proof. Merge-readiness now includes regulatory readiness.

ephor:governance · PASSED domain/evidence.py
04

Handoffs carry chain hashes

WeftMark handoffs are portable replay packages for handing work between agents. They carry the branch, base SHA, active scopes, failing tests, and next action — everything the receiving agent needs to continue. Adding the Ephor chain hash to the handoff record gives the receiving agent cryptographic proof that the prior session's actions were governed. Cross-session audit provenance becomes portable: the new agent can verify the handoff's lineage before operating, and Ephor can link its own audit entries across sessions using the same chain reference. Multi-agent work leaves a single continuous audit trail, not one per session.

handoff.chain_hash cross-session provenance
05

Evidence pack enriched with WeftMark lineage

Ephor's evidence pack export already bundles audit entries, policy snapshots, HITL decisions, and chain integrity proof for a conformity assessment body. Enriching it with WeftMark lineage — the Change Set graph, declared scope, evidence state, and reviewer decisions — produces a complete picture for Art 11 and Art 47 documentation: not just what the AI did, but what it was authorised to do, what engineering evidence existed at time of action, and which humans reviewed it. The conformity bundle becomes genuinely self-contained rather than pointing auditors to a separate engineering system they cannot interpret.

Art 11 · Art 47 evidence pack export

Also

Sylvae closes the execution layer

Sylvae

Sylvae is the portable skill runner: load a SKILL.md skill and run it against Ollama, Claude Code, Codex, OpenCode, or the Anthropic API — with every run logged as a durable evidence record. Sylvae's Phase 2 design already calls out a run_id as "the natural join key for future WeftMark ingestion."

With Ephor in the loop, each Sylvae run_id becomes an Ephor entry_id — every skill delegation is captured, policy-gated when the skill touches a high-risk scope, and appended to the immutable audit chain. The LLM judge calls in Sylvae's Phase 2 routing system are themselves AI decisions of exactly the kind Ephor was built to audit: consequential, automated, and in need of a record.


Roadmap · Rough

What needs to be built

These are the pieces. None of them are large in isolation — the biggest is agreeing on the shared protocol.

WeftMark

adapters/ephor.py — HTTP adapter for Ephor's /capture and /oversight API; lifecycle event fan-out on Change Set state transitions
ephor:governance evidence kind in domain/evidence.py; evidence policy that can require it as a merge precondition
Scope-to-risk mapping config: declare which semantic scope prefixes (contract:, security:) elevate Ephor risk level
chain_hash field on the Handoff domain model; include in the serialised handoff record and CLI output

Ephor

WeftMark session type in the agent taxonomy — agent_class: WeftMarkWorker with sub-type routing by backend (Codex, OpenCode, Ollama, etc.)
changeset_id as a first-class field on audit entries, indexed for lookup and cross-reference
WeftMark lineage export in the evidence pack: Change Set graph, declared scopes, evidence state, and reviewer decisions as a structured JSON section

Shared Protocol

Scope-risk mapping convention: agree on which WeftMark scope prefixes trigger which Ephor risk level; version and publish as a shared config schema
Agent class taxonomy: how WeftMark actors appear in the Ephor audit chain; how backend (Ollama vs Codex vs Claude) is recorded
Cross-reference standard: WeftMark evidence records cite Ephor chain hashes; Ephor entries cite WeftMark Change Set IDs. Bi-directional linking, stable across versions