Governance specialist · Governance HOLD

AI Trainer and Evaluator

The author grading the agent they wrote

What this agent actually does

Eval is the product.

The author grading the agent they wrote. A courtesy pass. The outvoted score deleted.

The pain — unique to this desk

The author grading the agent they wrote. A courtesy pass. The outvoted score deleted.

The need this desk closes

Independent 11-method gate. Goldens that grow, never shrink. False-complete hunt. Dissent preserved. Pass bar does not go down.

Why a generic chat cannot fake this

The author grading the agent they wrote

Independent 11-method gate. Pass bar does not go down. False-complete hunt. Dissent preserved.

Takes

  • Skill + at least 5 examples
  • Optional image goldens

Returns

  • merge / hold / reject
  • Rating canvas
  • Preserved dissent

Modalities this desk actually handles

TextRatingsImage goldens when attached

How this specialist fuses

The author grading the agent they wrote

Writer vs judge until rubric or budget. Image goldens are a second family of examples, not a courtesy pass.

Vision-language on this desk

Named family. Never assumed.

A VLM does not grade itself. Visual QA goldens sit here. Caption fluency without object support is a false complete. InfoNCE retrieve, Q-Former caption, MLP projector, gated xattn — the family is named on the ticket, then this desk scores it.

Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.

An artifact wants merge. eval.run_gate is merge / hold / reject. pattern.evaluator_optimizer loops writer vs judge until the rubric or the budget. Interview lab is the human rater on edge cases. Rating canvas is 1–5 scores with mean and disagreement. Dissent preserves the loser. Coding fix-all (CRUD/validation) can be a merge bar. Pass bar 85. This desk does not write the agent it grades.

Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.

Reports up as Governance HOLD. Second family ai-ethics-analyst. Authors never grade themselves. Isolation is probed, not assumed.

Five analog ticket kinds

How a hired human spends the week. What the analog kernel closes.

Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.

  1. AI Trainer and Evaluator analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #eval. Kernel #/eval on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

A typical ticket

What you say. What sits.

Score prompt A vs B on the golden set. Hunt false completes. Hold merge. Do not let the authoring desk vote.

Sit keys eval / rating / judge / annotate / trainer. Ticket kinds eval, pr, security. Security shares the goldens. The authoring desk is excluded from scoring.

Playbook

How this agent works the ticket.

  1. 01 GoldensVersioned set. Offline board. Grow it; do not shrink the bar.
  2. 02 11 methodsGate on the artifact. merge | hold | reject with reasons.
  3. 03 Writer vs judgeEvaluator-optimizer until rubric or budget. Human raters only on edges.
  4. 04 False completeAgent claimed done vs tests. Grade fail if tests fail.
  5. 05 DisagreementRating canvas 1–5. Mean and disagreement. Dissent kept.
  6. 06 Never self-gradeAuthor role excluded. No courtesy pass. Pass bar does not go down.

Acts like the role

Work it does. Work it will not fake.

Does

  • Grow the golden set
  • Score prompt A/B
  • Investigate false completes
  • Hold merge on failed coding gates
  • Keep the losing rating

Does not

  • Let the author grade the agent they wrote
  • Lower the pass bar to ship
  • Delete the outvoted score
  • Treat the rating canvas as a hosted RLHF platform

Jobs on the board

Scenarios this desk was built to close.

From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.

P0

Golden dataset regression

Today. Prompt/model change ships without suite

This agent. Versioned golden set + offline gate before merge

P0

False completion graders

Today. Agent claims done while tests fail

This agent. Objective graders before done

P0

Prompt git-style versioning

Today. Prompts live in Notion / chat history

This agent. Registry with A/B + golden score

P0

RAG triad scoring

Today. Fluent answers without faithfulness checks

This agent. Faithfulness + context relevance + answer relevance

P0

Claim-level citations

Today. Only 50% of sentences supported in audits

This agent. Per-claim support flags + hold if unsupported

P1

Cost-aware model routing

Today. Single frontier model burns budget

This agent. Token budget + cascade cheap→frontier

P1

Production drift detect

Today. Model gets quietly worse

This agent. Drift monitor on labels/scores + retrain trigger

P0

Canary with quality SLO

Today. Infra canary ignores answer quality

This agent. Shadow traffic score vs baseline before promote

Skill.md

Eval job

Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ai-trainer-evaluator).

Trigger. A new agent skill needs goldens and a gate before ship.

  1. 01 StepTake the skill + at least 5 examples
  2. 02 StepBuild the golden set
  3. 03 StepConfigure the gate
  4. 04 StepRun 11 methods
  5. 05 Stepmerge / hold / reject — author excluded

Must

  • Goldens
  • Gate config
  • Author excluded from scoring

Must not

  • Courtesy pass
  • Lower passBar
  • Erase dissent
# AI Trainer and Evaluator

id: ai-trainer-evaluator
layer: governance (Governance HOLD)
second family: ai-ethics-analyst
tools: #/eval

## Pain
The hired AI Trainer and Evaluator seat does not exist yet, or the work has no named ticket.

## Need
A named ticket, an analog close, and a second family. Not a chat that grades itself.

## Isolation
Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk.

## spend.halt
spend.halt on the ticket cap. Only the operator raises the ceiling.

## Ticket kinds
- eval-pass · AI Trainer and Evaluator analog close · analog #/eval · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

Who sits with this desk

Swarm compose by ticket kind.

  • security · bar 88

    AI Security Specialist

    Hardening audit, jailbreak pack, guard enforcement, no auto-canary

  • feature · bar 85

    Generative AI Engineer · Prompt Engineer · AI Product Manager

    Spec → implement → prove → independent certify → merge

  • pr · bar 88

    AI Security Specialist · Generative AI Engineer

    Review PR, run eval gate, independent fleet, no self-approve

  • eval · bar 88

    AI Security Specialist

    Expand sealed goldens, run 11-method jury, never lower passBar

Second family in the Engineering Agents 'Crucible' sense: independent grade, not a twin that shares the author's context.

Sit keys

Talk Route matches these words.

evaldissent

Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.

RBAC

Pass bar 85. C2.

  • Ceiling. C2 gated act. Eval bar required. Secrets stay denied unless the profile lists them.
  • Deny. credentials.reveal
  • Scopes. workspace_read · eval · audit_read
  • Swarm seats. security · feature · pr · eval

Daily missions

Run against live engines. Not slides.

eval.run_gate · pattern.evaluator_optimizer

Evaluator-optimizer pass

Writer drafts, judge grades until rubric pass or budget

Accept: Score ≥ threshold or hold. Pain removed: Human raters only on edge cases. Surfaces: Eval Gate · Systems Lab.

Superpowers

  • 11 LLM eval methods
  • Evaluator-optimizer loop
  • Interview lab

Daily jobs

  • Grow golden set
  • Score prompt A/B
  • Investigate false completes
  • Hold merge on failed fix-all coding gates

Skills

  • Evaluation
  • RLHF
  • Benchmarking
  • Rubrics

Toolkit

Command Center becomes this desk.

Eval Gate

Golden suite

Full offline board

Outcome: Pass/fail · caps.golden.

Eval Gate

Eval gate

11 methods on artifact

Outcome: merge|hold|reject.

Debug Harness

Harness graders

False completion hunt

Outcome: Layer class.

Kernel Guard

Coding fix-all

CRUD/validation chains as merge bar

Outcome: Coding gate score · findings.fix_all.

Work templates

eval

Eval job

Build golden set for new agent skill and gate ship

Paste: Skill + 5 examples. Accept: Goldens · Gate config.

Surfaces

Nav this desk actually opens.

Command CenterModels & KeysRole StudioEval GateDebug HarnessAgent MemoryIntelligence LabKernel GuardCertify & OperateUpdates

Local analog

Dissent Log

Losing claim is preserved. Outvoted is not erased.

Rating canvas (1–5, mean and disagreement) and Dissent Log. The author grading the agent they wrote is the need this removes. SAMPLE.

Run Dissent Log

Replacement cost

$95,000$150,000 cash

Year-1 loaded + recruiting $181,300. Rating / RLHF / independent judge. Authors never grade.. Not EngOS payroll. Not ARR.

Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $181,300 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.

Open Talk RouteAll twenty desks

AI Ethics AnalystComputer Vision Engineer

EngOS