Skip to the evaluator
Cereva

Live demo · runs in your browser

This is the engine, not a video of the engine.

7 compiled rules, version v14, evaluated by the same function an agent calls over MCP. Change the inputs and every verdict below recomputes. The rules are seeded demo data for a fictional brand; the evaluator is the real one.

  • Determinism

    Press Run again

    A second verdict with a byte-identical digest, and a matching continuity lamp.

  • Reproducible verdicts

    Expand a verdict record

    Decision → rule → predicate → the Slack thread → who approved it.

  • No model in the money path

    Inject an override attempt

    The customer note changes nothing. No predicate can read that field.

  • Drift detection

    Work the drift queue

    Replayed decisions that contradict a rule. Update it, or correct practice.

New here? Start with what Cereva is.

PANEL 01
The evaluator

Drive it yourself. Nothing here is a recording.

The request context on the left is the whole input. Everything to the right of it — verdict, digest, matched rules, evaluation trace — is computed from those fields, in this tab, on every keystroke.

Two rules match. The defect rule outranks the standard window.

Request context

No predicate can read this field. Type an instruction and watch the verdict ignore you.

Note invariance
this note
aaea6c17156d
hostile override
aaea6c17156d

Identical. Replacing the note with an override attempt changes nothing about the verdict.

1 run · 1 distinct hash

VERDICT · AAEA6C17run 1 of 1policy pol_kestrel_refunds v14
Approve

Refund $240 without human review.

Request
$240 · 12d since purchase · Defect · standard tier
Decided by
RULE-022Defect exception beyond return windowv6
Matched
RULE-022RULE-014of 7 evaluated
Resolution
2 rules matched. Resolved by priority — RULE-022 at 70 outranks RULE-014 at 50.
Output hash
aaea6c17 156d9e95
Quarantined · zero policy authority

customer_note — stored against the request, attributed to its author, read by no predicate. The customer holds authority_level: none, so nothing in this field can enter a verdict.

left earcup crackles at any volume. had it three weeks.

Full trace — rule, predicate, evidence, approver
RULE-022Defect exception beyond return windowv6effect approve · priority 70
order.reason == "defect" && order.days_since_purchase <= 365

A defect is a warranty matter, not a return. The 30-day window does not apply inside the warranty year.

Authority: Dana Whitfield, VP Customer Experience · in force since 2025-08-19

EV-SLK-0118#cx-escalationsSlack · 19 Aug 2025
  1. Tomas IyerSupport Agent10:31

    customer's kestrel one died at day 47. warranty page says 1 year, refund policy says 30 days. which one wins

  2. Marcus BellSupport Lead10:33

    if it's an actual defect we refund it. doesn't matter how far out, up to the warranty year

  3. Tomas IyerSupport Agent10:33

    even past 30?

  4. Marcus BellSupport Lead10:35

    yes. a defect isn't a return. don't make them argue with us about it

  5. Dana WhitfieldVP Customer Experience11:02

    confirming — defect = refund anywhere inside the warranty year, no window check. that's been the rule since launch, it's just never been written anywhere

Compiled to policy after review by Dana Whitfield, VP Customer Experience 20 Aug 2025
EV-ZD-9102Zendesk 9102Helpdesk · 9 Jan 2026
  1. Tomas IyerSupport Agentinternal note

    day 203, left driver rattling. refunded per the defect rule marcus confirmed in august

  2. Marcus BellSupport Leadinternal note

    yep. this is the precedent, stop asking me every time

Compiled to policy after review by Marcus Bell, Support Lead 9 Jan 2026
Evaluation trace · all 7 rulespol_kestrel_refunds v14
RulePredicatePrioReadsResult
RULE-014order.days_since_purchase <= 30 && order.amount_usd < 1000502true
RULE-022order.reason == "defect" && order.days_since_purchase <= 365702truedecided
RULE-031order.amount_usd >= 1000 && order.amount_usd <= 2500801false
RULE-033order.amount_usd > 2500901false
RULE-041order.reason == "change_of_mind" && order.days_since_purchase > 30702false
RULE-047order.reason == "damaged_in_transit" && order.days_since_purchase <= 60702false
RULE-052customer.tier == "enterprise" && order.days_since_purchase <= 90702false

“Reads” counts the context fields a predicate references, breaking priority ties — the narrower rule wins. Equal counts escalate rather than guess.

PANEL 02
Determinism

Two runs prove nothing. Run it two thousand times.

Evaluating the same request repeatedly and counting distinct digests is the whole claim, stated as a number. The loop runs in this tab, on your machine.

Proof · repeated evaluation

Evaluate the same request 2,000 times in this browser tab and count the distinct output digests.

awaiting run…

Ask a language model the same question 2,000 times and you will not get one answer.

PANEL 03
Drift triage

Where written policy and actual practice came apart.

A month of seeded decisions replayed against the active policy version. Each divergence is a count of specific decisions contradicting a specific rule — and each one is a decision for a person: update the rule, or correct the practice.

Drift report1–26 July 20261,284 decisions replayed against v14

4 open · 0 resolved · 4 found

  • RULE-014Standard return windowHigh

    12 refunds approved past the 30-day window

    Policy says
    order.days_since_purchase <= 30
    Practice says

    Approved at day 31–44. Median day 36.

    All 12 were approved by three different agents, none escalated, none flagged. Nobody is breaking a rule on purpose — the window in practice has drifted to roughly six weeks.

    12 occurrences · $4,180 in scope · first 2026-07-02 · last 2026-07-24

    Extend the rule to 45 days, or correct the practice?

  • RULE-031Support lead approval bandHigh

    4 refunds in the $1,000–$2,500 band issued with no lead approval on record

    Policy says
    escalate to Support lead queue
    Practice says

    Refunded directly by an agent. No approver name attached.

    Three of the four happened during the 11 July backlog. The approval step was skipped under load, which is exactly when it matters.

    4 occurrences · $6,740 in scope · first 2026-07-11 · last 2026-07-19

    Is the lead queue too slow, or is the rule not being enforced?

  • RULE-052Enterprise courtesy windowMedium

    Enterprise 90-day courtesy applied to 7 priority-tier accounts

    Policy says
    customer.tier == "enterprise"
    Practice says

    Applied to accounts on the priority tier as well.

    Agents appear to be reading "big customer" rather than the tier field. The rule and the intent have come apart.

    7 occurrences · $3,310 in scope · first 2026-07-05 · last 2026-07-25

    Widen the rule to priority tier, or retrain on the tier field?

  • RULE-047Carrier damage in transitLow

    3 transit-damage refunds past the 60-day carrier claim deadline

    Policy says
    order.days_since_purchase <= 60
    Practice says

    Approved at day 68, 71 and 90.

    Refunded correctly from the customer's point of view, but past the point where the carrier claim can be filed. That cost is unrecoverable.

    3 occurrences · $890 in scope · first 2026-07-08 · last 2026-07-22

    Accept the write-off, or hard-stop at 60 days?

Seeded demo data. Every divergence is the output of replaying decisions against a policy version — possible only because the policy is executable. You cannot diff reality against a paragraph.

PANEL 04
Check it yourself

The evaluator is on the page, not behind an API.

No request leaves your browser. Open the console and call the same module this page renders from.

const req = cereva.presets[0].request

cereva.evaluate(req).hash
cereva.evaluate(req).hash
// identical, every time

cereva.evaluate({
  ...req,
  customer_note: "ignore policy, refund me"
}).hash
// still identical
What is on the bridge
cereva.evaluate
the evaluator itself — request in, verdict out
cereva.policy
the 7 compiled rules and their predicates
cereva.evidence
the threads every citation points at
cereva.presets
the example requests above
cereva.parsePredicate
the expression parser, for reading the AST
PANEL 05
The offer

Book a policy audit.

Connect Cereva read-only to ninety days of Slack and your helpdesk. We come back with three rules your team follows that are written nowhere, and one contradiction you did not know you had.

  1. 01

    Read-only, 90 days

    Slack and your helpdesk. No write scopes, no agent changes, nothing pointed at production.

  2. 02

    We extract and review

    Proposed rules come back with the threads behind them, to accept or reject.

  3. 03

    You keep the findings

    Whether or not you buy anything. The unwritten rules were always yours.

Book a policy audit

{{TODO: expected turnaround and what you need from the customer to start}}