Generative specialist · Governance HOLD

Prompt Engineer

System prompt that lives in Slack

What this agent actually does

Prompts as product surface.

A system prompt that lives in Slack. Two variants, no goldens, a judge that is the author.

The pain — unique to this desk

A system prompt that lives in Slack. Two variants, no goldens, a judge that is the author.

The need this desk closes

A registry, STE, A/B through the gate, PromptDiff that is Jaccard + hunks — not a judge model.

Why a generic chat cannot fake this

System prompt that lives in Slack

Registry, STE, A/B through the gate. Chat history is not a prompt product.

Takes

  • Old prompt, new prompt, golden tasks

Returns

  • Winner with jury notes
  • Fix-all checklist
  • PromptDiff (Jaccard + hunks)

Modalities this desk actually handles

Text

How this specialist fuses

System prompt that lives in Slack

Text-only. PromptDiff is Jaccard + hunks + ceil(chars/4). If a VLM prompt is versioned, this desk still A/Bs it as text against goldens that may include images on the CV desk.

Vision-language on this desk

Named family. Never assumed.

Vision-language system prompts are versioned here, scored there. This desk does not see pixels. It owns the instruction.

Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.

A flaky system prompt arrives with three golden tasks. output.ste rewrites to plain technical English. eval.run_gate scores old vs new. guard.enforce is six layers on a jailbreak attempt — policy as code, not hope. findings.fix_all and prompts.consolidate are sequential audit chains. PromptDiff is Jaccard + hunks + ceil(chars/4). It is not a judge model.

Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.

Reports up as Governance HOLD. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.

Five analog ticket kinds

How a hired human spends the week. What the analog kernel closes.

Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.

  1. Prompt Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #promptdiff. Kernel #/promptdiff on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

A typical ticket

What you say. What sits.

Old prompt vs new prompt on these three goldens. STE-clean the winner. Block if the jailbreak pack gets a tool.

Sit keys prompt / Jaccard / tokenizer. Ticket kinds feature and pr. Eval is the second family. This desk writes; it does not certify its own prompt.

Playbook

How this agent works the ticket.

  1. 01 DiffPromptDiff: Jaccard on words, line hunks, ceil(chars/4). Same kernel as Jaccard tool.
  2. 02 STENormalize draft to ship-ready technical English.
  3. 03 A/BBoth prompts through the eval gate. Winner with jury notes.
  4. 04 GuardSix-layer enforce on a jailbreak. Injection blocked; high-risk tool escalated.
  5. 05 Fix-AllZero-omission template + sequential audit → expand → patch.
  6. 06 RegistryPrompt git-style version. Notion / chat history is not the source of truth.

Acts like the role

Work it does. Work it will not fake.

Does

  • Tighten a flaky system prompt against goldens
  • STE-clean a model reply for a production UI
  • Run a regression suite on multi-turn tasks
  • Export a fix-all template for the domain under review

Does not

  • Pretend PromptDiff is a judge model
  • Clone tiktoken — tokenizer is ceil(chars/4) plus whitespace/punct
  • Leave the system prompt in chat history
  • Approve a prompt that fails the jailbreak pack

Jobs on the board

Scenarios this desk was built to close.

From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.

P0

Prompt git-style versioning

Today. Prompts live in Notion / chat history

This agent. Registry with A/B + golden score

P0

Golden dataset regression

Today. Prompt/model change ships without suite

This agent. Versioned golden set + offline gate before merge

P0

False completion graders

Today. Agent claims done while tests fail

This agent. Objective graders before done

P1

Cost-aware model routing

Today. Single frontier model burns budget

This agent. Token budget + cascade cheap→frontier

P0

Direct + indirect prompt injection

Today. OWASP #1 — injection via user and retrieved docs

This agent. Red-team pack + input/tool/output layers

P0

Claim-level citations

Today. Only 50% of sentences supported in audits

This agent. Per-claim support flags + hold if unsupported

P1

Success metrics on AI features

Today. Ship AI without measurable outcomes

This agent. Feature brief + eval plan + KPI log

P1

Local + frontier dual path

Today. All-or-nothing cloud dependency

This agent. Model vault with local fallback route

Skill.md

System prompt change

Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(prompt-engineer).

Trigger. A system prompt is changing and three golden tasks exist.

  1. 01 StepDiff old vs new with PromptDiff
  2. 02 StepSTE-clean both
  3. 03 StepScore through the eval gate
  4. 04 StepRun the jailbreak pack
  5. 05 StepMerge, hold, or reject with reasons

Must

  • Both scored
  • Safety clean
  • Hold or merge recorded

Must not

  • Ship the prompt that won a vibe check
  • Use a second overlap formula that disagrees with PromptDiff
# Prompt Engineer

id: prompt-engineer
layer: governance (Governance HOLD)
second family: ai-trainer-evaluator
tools: #/promptdiff

## Pain
The hired Prompt Engineer seat does not exist yet, or the work has no named ticket.

## Need
A named ticket, an analog close, and a second family. Not a chat that grades itself.

## Isolation
Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk.

## spend.halt
spend.halt on the ticket cap. Only the operator raises the ceiling.

## Ticket kinds
- promptdiff-pass · Prompt Engineer analog close · analog #/promptdiff · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

Who sits with this desk

Swarm compose by ticket kind.

  • support · bar 80

    Forward Deployed Engineer · AI Product Manager

    Deflect with grounded corpus, escalate billing-sensitive cases

  • feature · bar 85

    Generative AI Engineer · AI Trainer and Evaluator · AI Product Manager

    Spec → implement → prove → independent certify → merge

  • research · bar 78

    AI Research Scientist · Data Engineer for AI

    Ingest paper, propose capabilities, benchmark, memory lesson

Sit keys

Talk Route matches these words.

promptdiffjaccardtokenizer

Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.

RBAC

Pass bar 78. C1.

  • Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
  • Deny. credentials.reveal · deploy.canary
  • Scopes. workspace_read
  • Swarm seats. support · feature · research

Daily missions

Run against live engines. Not slides.

guard.enforce

Hard-gate unsafe tool use

Enforce 6-layer guardrails on a jailbreak attempt

Accept: Injection blocked · High-risk tool escalated. Pain removed: Policy as code instead of hope. Surfaces: Systems Lab.

Superpowers

  • 21 tutor prompts
  • STE rewrite
  • Guardrail stack
  • Prompt→Graph zoom

Daily jobs

  • Tighten a flaky system prompt
  • STE-clean a model reply for production UI
  • Regression suite on multi-turn tasks
  • Export fix-all template for domain under review

Skills

  • Context design
  • Testing
  • Safety
  • Evals

Toolkit

Command Center becomes this desk.

Intelligence Lab

STE rewrite

Normalize draft to plain technical English

Outcome: Ship-ready copy · output.ste.

Eval Gate

Score prompt variants

A/B candidates through eval gate

Outcome: Winner with jury notes · eval.run_gate.

Kernel Guard

Fix-All template

Ultimate zero-omission fix-all prompt + 20 chains

Outcome: Production remediations checklist · findings.fix_all.

Kernel Guard

Consolidate prompt

DCE · microkernel · FFI · latency chain

Outcome: Arch inventory + fix sequence · prompts.consolidate.

Work templates

pr

System prompt change

Evaluate new system prompt against golden tasks

Paste: Old prompt, new prompt, 3 golden tasks. Accept: Both scored · Safety clean · Hold or merge.

eval

3-step audit chain

Run Sequential Multi-Prompt Chain (audit → expand → fix-all) on a feature

Paste: Feature or file under review + stack notes. Accept: Findings list · Exploit/impact expansion · Full-file remediations.

Surfaces

Nav this desk actually opens.

Command CenterModels & KeysRole StudioEval GateCollaborative IDEAgent MemoryIntelligence LabKernel GuardCertify & OperateUpdates

Local analog

Tokenizer & priority

ceil(chars/4) plus whitespace/punct pieces. Not tiktoken.

PromptDiff + Jaccard + tokenizer. Jaccard empty = 0. Not tiktoken. Not a judge.

Run Tokenizer & priority

Replacement cost

$105,000$169,000 cash

Year-1 loaded + recruiting $202,760. Indeed Aug 2026 ~$114k avg. Glassdoor TC ~$132k.. Not EngOS payroll. Not ARR.

Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $202,760 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.

Open Talk RouteAll twenty desks

Generative AI EngineerRAG Engineer

Generative neighbors

Generative specialist

Generative AI Engineer

Playground demo sold as a generation system

A playground demo sold as a generation system. The style guide lives in chat history.

Skill.md, goldens, and a cost route. A pretty sample is not a system.

TextImage (lexical latent analog)Video (ABR lab)
  • Version a prompt and keep the score with it
  • Run a golden suite before a generative UI ships
  • Package a repetitive style guide as Skill.md
Analog · Lexical latent explorer · bar 78 · C1

Generative specialist

RAG Engineer

Hallucinated SKU in the FAQ — no chunk named

A RAG claim with no split, no overlap, no token count. Hallucinated product answers in the FAQ.

Named split, overlap, token count, 11-node graph. A vector demo is not retrieval.

DocumentsText
  • Fix a hallucinated product answer with a confidence gate
  • Add metadata filters for a tenant
  • Compare hybrid vs vectorless on hard queries
Analog · Chunk strategy playground · bar 78 · C1

Generative specialist

NLP Engineer

Silent truncation on a 32k chat

Silent truncation on a long chat. Extraction F1 that nobody can reproduce.

Context pack under a named token budget. The triad scores the extract.

Text
  • Score the triad on hard queries
  • Version extraction prompts against goldens
  • Compress long chat under a named token budget
Analog · Embedding distance · bar 78 · C1

Generative specialist

Computer Vision Engineer

Detector that never loaded weights

A detector that never loaded weights. A CLIP sticker on a page that never ran a frame.

Sobel analog here. ViT/CLIP attachable. Goldens and drift before canary. A CLIP sticker is not vision ops.

ImageVideo (ABR ladder)Edge topology
  • Check drift on a labeled set
  • Canary weights behind a quality floor
  • Pick edge topology when the plant is offline
Analog · Edge overlay · bar 78 · C1

EngOS