The pain — unique to this desk
Generative specialist · Governance HOLD
Prompt Engineer
System prompt that lives in Slack
What this agent actually does
Prompts as product surface.
A system prompt that lives in Slack. Two variants, no goldens, a judge that is the author.
The need this desk closes
A registry, STE, A/B through the gate, PromptDiff that is Jaccard + hunks — not a judge model.
Why a generic chat cannot fake this
System prompt that lives in Slack
Registry, STE, A/B through the gate. Chat history is not a prompt product.
Takes
- Old prompt, new prompt, golden tasks
Returns
- Winner with jury notes
- Fix-all checklist
- PromptDiff (Jaccard + hunks)
Modalities this desk actually handles
How this specialist fuses
System prompt that lives in Slack
Text-only. PromptDiff is Jaccard + hunks + ceil(chars/4). If a VLM prompt is versioned, this desk still A/Bs it as text against goldens that may include images on the CV desk.
Vision-language on this desk
Named family. Never assumed.
Vision-language system prompts are versioned here, scored there. This desk does not see pixels. It owns the instruction.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
A flaky system prompt arrives with three golden tasks. output.ste rewrites to plain technical English. eval.run_gate scores old vs new. guard.enforce is six layers on a jailbreak attempt — policy as code, not hope. findings.fix_all and prompts.consolidate are sequential audit chains. PromptDiff is Jaccard + hunks + ceil(chars/4). It is not a judge model.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as Governance HOLD. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- Prompt Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #promptdiff. Kernel #/promptdiff on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
Old prompt vs new prompt on these three goldens. STE-clean the winner. Block if the jailbreak pack gets a tool.
Sit keys prompt / Jaccard / tokenizer. Ticket kinds feature and pr. Eval is the second family. This desk writes; it does not certify its own prompt.
Playbook
How this agent works the ticket.
- 01 DiffPromptDiff: Jaccard on words, line hunks, ceil(chars/4). Same kernel as Jaccard tool.
- 02 STENormalize draft to ship-ready technical English.
- 03 A/BBoth prompts through the eval gate. Winner with jury notes.
- 04 GuardSix-layer enforce on a jailbreak. Injection blocked; high-risk tool escalated.
- 05 Fix-AllZero-omission template + sequential audit → expand → patch.
- 06 RegistryPrompt git-style version. Notion / chat history is not the source of truth.
Acts like the role
Work it does. Work it will not fake.
Does
- Tighten a flaky system prompt against goldens
- STE-clean a model reply for a production UI
- Run a regression suite on multi-turn tasks
- Export a fix-all template for the domain under review
Does not
- Pretend PromptDiff is a judge model
- Clone tiktoken — tokenizer is ceil(chars/4) plus whitespace/punct
- Leave the system prompt in chat history
- Approve a prompt that fails the jailbreak pack
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P0
False completion graders
Today. Agent claims done while tests fail
This agent. Objective graders before done
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P0
Direct + indirect prompt injection
Today. OWASP #1 — injection via user and retrieved docs
This agent. Red-team pack + input/tool/output layers
P0
Claim-level citations
Today. Only 50% of sentences supported in audits
This agent. Per-claim support flags + hold if unsupported
P1
Success metrics on AI features
Today. Ship AI without measurable outcomes
This agent. Feature brief + eval plan + KPI log
P1
Local + frontier dual path
Today. All-or-nothing cloud dependency
This agent. Model vault with local fallback route
Skill.md
System prompt change
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(prompt-engineer).
Trigger. A system prompt is changing and three golden tasks exist.
- 01 StepDiff old vs new with PromptDiff
- 02 StepSTE-clean both
- 03 StepScore through the eval gate
- 04 StepRun the jailbreak pack
- 05 StepMerge, hold, or reject with reasons
Must
- Both scored
- Safety clean
- Hold or merge recorded
Must not
- Ship the prompt that won a vibe check
- Use a second overlap formula that disagrees with PromptDiff
# Prompt Engineer id: prompt-engineer layer: governance (Governance HOLD) second family: ai-trainer-evaluator tools: #/promptdiff ## Pain The hired Prompt Engineer seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - promptdiff-pass · Prompt Engineer analog close · analog #/promptdiff · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
support · bar 80
Forward Deployed Engineer · AI Product Manager
Deflect with grounded corpus, escalate billing-sensitive cases
feature · bar 85
Generative AI Engineer · AI Trainer and Evaluator · AI Product Manager
Spec → implement → prove → independent certify → merge
research · bar 78
AI Research Scientist · Data Engineer for AI
Ingest paper, propose capabilities, benchmark, memory lesson
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 78. C1.
- Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
- Deny. credentials.reveal · deploy.canary
- Scopes. workspace_read
- Swarm seats. support · feature · research
Daily missions
Run against live engines. Not slides.
guard.enforce
Hard-gate unsafe tool use
Enforce 6-layer guardrails on a jailbreak attempt
Accept: Injection blocked · High-risk tool escalated. Pain removed: Policy as code instead of hope. Surfaces: Systems Lab.
Superpowers
- 21 tutor prompts
- STE rewrite
- Guardrail stack
- Prompt→Graph zoom
Daily jobs
- Tighten a flaky system prompt
- STE-clean a model reply for production UI
- Regression suite on multi-turn tasks
- Export fix-all template for domain under review
Skills
- Context design
- Testing
- Safety
- Evals
Toolkit
Command Center becomes this desk.
Intelligence Lab
STE rewrite
Normalize draft to plain technical English
Outcome: Ship-ready copy · output.ste.
Eval Gate
Score prompt variants
A/B candidates through eval gate
Outcome: Winner with jury notes · eval.run_gate.
Kernel Guard
Fix-All template
Ultimate zero-omission fix-all prompt + 20 chains
Outcome: Production remediations checklist · findings.fix_all.
Kernel Guard
Consolidate prompt
DCE · microkernel · FFI · latency chain
Outcome: Arch inventory + fix sequence · prompts.consolidate.
Work templates
pr
System prompt change
Evaluate new system prompt against golden tasks
Paste: Old prompt, new prompt, 3 golden tasks. Accept: Both scored · Safety clean · Hold or merge.
eval
3-step audit chain
Run Sequential Multi-Prompt Chain (audit → expand → fix-all) on a feature
Paste: Feature or file under review + stack notes. Accept: Findings list · Exploit/impact expansion · Full-file remediations.
Surfaces
Nav this desk actually opens.
Local analog
Tokenizer & priority
ceil(chars/4) plus whitespace/punct pieces. Not tiktoken.
PromptDiff + Jaccard + tokenizer. Jaccard empty = 0. Not tiktoken. Not a judge.
Replacement cost
$105,000–$169,000 cash
Year-1 loaded + recruiting $202,760. Indeed Aug 2026 ~$114k avg. Glassdoor TC ~$132k.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $202,760 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.
Generative neighbors
Generative specialist
Generative AI Engineer
Playground demo sold as a generation system
A playground demo sold as a generation system. The style guide lives in chat history.
Skill.md, goldens, and a cost route. A pretty sample is not a system.
- Version a prompt and keep the score with it
- Run a golden suite before a generative UI ships
- Package a repetitive style guide as Skill.md
Generative specialist
RAG Engineer
Hallucinated SKU in the FAQ — no chunk named
A RAG claim with no split, no overlap, no token count. Hallucinated product answers in the FAQ.
Named split, overlap, token count, 11-node graph. A vector demo is not retrieval.
- Fix a hallucinated product answer with a confidence gate
- Add metadata filters for a tenant
- Compare hybrid vs vectorless on hard queries
Generative specialist
NLP Engineer
Silent truncation on a 32k chat
Silent truncation on a long chat. Extraction F1 that nobody can reproduce.
Context pack under a named token budget. The triad scores the extract.
- Score the triad on hard queries
- Version extraction prompts against goldens
- Compress long chat under a named token budget
Generative specialist
Computer Vision Engineer
Detector that never loaded weights
A detector that never loaded weights. A CLIP sticker on a page that never ran a frame.
Sobel analog here. ViT/CLIP attachable. Goldens and drift before canary. A CLIP sticker is not vision ops.
- Check drift on a labeled set
- Canary weights behind a quality floor
- Pick edge topology when the plant is offline