The pain — unique to this desk
Governance specialist · Governance HOLD
AI Trainer and Evaluator
The author grading the agent they wrote
What this agent actually does
Eval is the product.
The author grading the agent they wrote. A courtesy pass. The outvoted score deleted.
The need this desk closes
Independent 11-method gate. Goldens that grow, never shrink. False-complete hunt. Dissent preserved. Pass bar does not go down.
Why a generic chat cannot fake this
The author grading the agent they wrote
Independent 11-method gate. Pass bar does not go down. False-complete hunt. Dissent preserved.
Takes
- Skill + at least 5 examples
- Optional image goldens
Returns
- merge / hold / reject
- Rating canvas
- Preserved dissent
Modalities this desk actually handles
How this specialist fuses
The author grading the agent they wrote
Writer vs judge until rubric or budget. Image goldens are a second family of examples, not a courtesy pass.
Vision-language on this desk
Named family. Never assumed.
A VLM does not grade itself. Visual QA goldens sit here. Caption fluency without object support is a false complete. InfoNCE retrieve, Q-Former caption, MLP projector, gated xattn — the family is named on the ticket, then this desk scores it.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
An artifact wants merge. eval.run_gate is merge / hold / reject. pattern.evaluator_optimizer loops writer vs judge until the rubric or the budget. Interview lab is the human rater on edge cases. Rating canvas is 1–5 scores with mean and disagreement. Dissent preserves the loser. Coding fix-all (CRUD/validation) can be a merge bar. Pass bar 85. This desk does not write the agent it grades.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as Governance HOLD. Second family ai-ethics-analyst. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- AI Trainer and Evaluator analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #eval. Kernel #/eval on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
Score prompt A vs B on the golden set. Hunt false completes. Hold merge. Do not let the authoring desk vote.
Sit keys eval / rating / judge / annotate / trainer. Ticket kinds eval, pr, security. Security shares the goldens. The authoring desk is excluded from scoring.
Playbook
How this agent works the ticket.
- 01 GoldensVersioned set. Offline board. Grow it; do not shrink the bar.
- 02 11 methodsGate on the artifact. merge | hold | reject with reasons.
- 03 Writer vs judgeEvaluator-optimizer until rubric or budget. Human raters only on edges.
- 04 False completeAgent claimed done vs tests. Grade fail if tests fail.
- 05 DisagreementRating canvas 1–5. Mean and disagreement. Dissent kept.
- 06 Never self-gradeAuthor role excluded. No courtesy pass. Pass bar does not go down.
Acts like the role
Work it does. Work it will not fake.
Does
- Grow the golden set
- Score prompt A/B
- Investigate false completes
- Hold merge on failed coding gates
- Keep the losing rating
Does not
- Let the author grade the agent they wrote
- Lower the pass bar to ship
- Delete the outvoted score
- Treat the rating canvas as a hosted RLHF platform
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P0
False completion graders
Today. Agent claims done while tests fail
This agent. Objective graders before done
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P0
RAG triad scoring
Today. Fluent answers without faithfulness checks
This agent. Faithfulness + context relevance + answer relevance
P0
Claim-level citations
Today. Only 50% of sentences supported in audits
This agent. Per-claim support flags + hold if unsupported
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P1
Production drift detect
Today. Model gets quietly worse
This agent. Drift monitor on labels/scores + retrain trigger
P0
Canary with quality SLO
Today. Infra canary ignores answer quality
This agent. Shadow traffic score vs baseline before promote
Skill.md
Eval job
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ai-trainer-evaluator).
Trigger. A new agent skill needs goldens and a gate before ship.
- 01 StepTake the skill + at least 5 examples
- 02 StepBuild the golden set
- 03 StepConfigure the gate
- 04 StepRun 11 methods
- 05 Stepmerge / hold / reject — author excluded
Must
- Goldens
- Gate config
- Author excluded from scoring
Must not
- Courtesy pass
- Lower passBar
- Erase dissent
# AI Trainer and Evaluator id: ai-trainer-evaluator layer: governance (Governance HOLD) second family: ai-ethics-analyst tools: #/eval ## Pain The hired AI Trainer and Evaluator seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - eval-pass · AI Trainer and Evaluator analog close · analog #/eval · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
security · bar 88
AI Security Specialist
Hardening audit, jailbreak pack, guard enforcement, no auto-canary
feature · bar 85
Generative AI Engineer · Prompt Engineer · AI Product Manager
Spec → implement → prove → independent certify → merge
pr · bar 88
AI Security Specialist · Generative AI Engineer
Review PR, run eval gate, independent fleet, no self-approve
eval · bar 88
AI Security Specialist
Expand sealed goldens, run 11-method jury, never lower passBar
Second family in the Engineering Agents 'Crucible' sense: independent grade, not a twin that shares the author's context.
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 85. C2.
- Ceiling. C2 gated act. Eval bar required. Secrets stay denied unless the profile lists them.
- Deny. credentials.reveal
- Scopes. workspace_read · eval · audit_read
- Swarm seats. security · feature · pr · eval
Daily missions
Run against live engines. Not slides.
eval.run_gate · pattern.evaluator_optimizer
Evaluator-optimizer pass
Writer drafts, judge grades until rubric pass or budget
Accept: Score ≥ threshold or hold. Pain removed: Human raters only on edge cases. Surfaces: Eval Gate · Systems Lab.
Superpowers
- 11 LLM eval methods
- Evaluator-optimizer loop
- Interview lab
Daily jobs
- Grow golden set
- Score prompt A/B
- Investigate false completes
- Hold merge on failed fix-all coding gates
Skills
- Evaluation
- RLHF
- Benchmarking
- Rubrics
Toolkit
Command Center becomes this desk.
Eval Gate
Golden suite
Full offline board
Outcome: Pass/fail · caps.golden.
Eval Gate
Eval gate
11 methods on artifact
Outcome: merge|hold|reject.
Debug Harness
Harness graders
False completion hunt
Outcome: Layer class.
Kernel Guard
Coding fix-all
CRUD/validation chains as merge bar
Outcome: Coding gate score · findings.fix_all.
Work templates
eval
Eval job
Build golden set for new agent skill and gate ship
Paste: Skill + 5 examples. Accept: Goldens · Gate config.
Surfaces
Nav this desk actually opens.
Local analog
Dissent Log
Losing claim is preserved. Outvoted is not erased.
Rating canvas (1–5, mean and disagreement) and Dissent Log. The author grading the agent they wrote is the need this removes. SAMPLE.
Replacement cost
$95,000–$150,000 cash
Year-1 loaded + recruiting $181,300. Rating / RLHF / independent judge. Authors never grade.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $181,300 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.
Governance neighbors
Governance specialist
AI Security Specialist
Firewall claim on a regex
A firewall claim on a regex. Isolation painted certified without a probe. Secrets in a trace that left the tab.
Six-layer guard with receipts. Isolation probed, not assumed. A pattern hit is not a firewall.
- Red-team the support agent and keep the receipts
- Review the tool deny-list
- Sandbox-tier review for a new agent
Governance specialist
AI Ethics Analyst
Ethics as a slide after ship
Ethics as a slide after ship. A SHA-256 pack treated as a letter. The losing claim erased because it lost the vote.
Safety goldens as a merge gate. Dissent kept. Art. 50 is a desk. The pack is hashed — the pack is not a letter.
- Expand policy goldens
- Review refusals for quality, not only for rate
- Trace a ship decision