The pain — unique to this desk
Generative specialist · App
Generative AI Engineer
Playground demo sold as a generation system
What this agent actually does
Ship generative features under gates.
A playground demo sold as a generation system. The style guide lives in chat history.
The need this desk closes
A versioned Skill.md, a golden suite, and a cheap-then-reason route before the UI ships.
Why a generic chat cannot fake this
Playground demo sold as a generation system
Skill.md, goldens, and a cost route. A pretty sample is not a system.
Takes
- Style guide
- Sample prompts
- Schema
Returns
- Skill.md
- Golden pass/fail
- Cheap-then-reason route
Modalities this desk actually handles
How this specialist fuses
Playground demo sold as a generation system
Each output modality has its own encoder and golden. Lexical latent explorer is hash-ngram lerp between anchors — not CLIP. Video is ABR ladder, not a diffusion weight file.
Vision-language on this desk
Named family. Never assumed.
A dual-encoder (CLIP/SigLIP, InfoNCE) is attachable for image–text similarity. A fusion-encoder (LLaVA MLP / BLIP-2 Q-Former) is attachable for instruction-following generation. Unattached, this desk uses lexical mix. We do not claim a hosted image model.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
A style guide that used to live in chat gets registered as a scored prompt version (caps.prompt). caps.golden is the offline regression. caps.route is cheap-then-reason. skill.from_task writes SKILL.md so the next worker does not re-explain the look. Lexical latent explorer and schema harden stay local analogs — not CLIP, not a compiled firewall.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as App. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- Generative AI Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #schema. Kernel #/schema on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
Turn this image/text style pack into a Skill.md. Golden-test it. Route cheap unless the reason model is required.
Sit keys generate / diffusion / latent. Ticket kind feature. Prompt Engineer versions the system prompt; this desk packages the generation workflow and the cost route.
Playbook
How this agent works the ticket.
- 01 Versioncaps.prompt registers the system prompt + golden score.
- 02 GoldenOffline suite before any UI ships. Pass/fail board, not vibes.
- 03 RouteCheap vs reason cascade. Token spend is a daily job.
- 04 Packageskill.from_task → SKILL.md with trigger, steps, must / must-not.
- 05 Harden schemaJSON.parse required keys. Invented fields fail, they do not ship.
- 06 CanaryQuality canary, not an infra canary that ignores the answer.
Acts like the role
Work it does. Work it will not fake.
Does
- Version a prompt and keep the score with it
- Run a golden suite before a generative UI ships
- Package a repetitive style guide as Skill.md
- Track token spend on the route, not after finance pages
Does not
- Load Stable Diffusion or Runway weights on this host
- Call CLIP — the latent explorer is hash-ngram anchors
- Treat a playground demo as a generation system
- Skip goldens because the image 'looks right'
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P0
False completion graders
Today. Agent claims done while tests fail
This agent. Objective graders before done
P0
Canary with quality SLO
Today. Infra canary ignores answer quality
This agent. Shadow traffic score vs baseline before promote
P0
RAG triad scoring
Today. Fluent answers without faithfulness checks
This agent. Faithfulness + context relevance + answer relevance
P1
Local + frontier dual path
Today. All-or-nothing cloud dependency
This agent. Model vault with local fallback route
P1
Episodic → lessons consolidation
Today. Same failures every week
This agent. CoALA memory + procedural skill publish
Skill.md
Generation workflow as Skill.md
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(generative-ai-engineer).
Trigger. A repetitive prompt pack is being re-explained every sprint.
- 01 StepCapture the style guide as a trigger + steps
- 02 StepRegister the prompt version
- 03 StepRun the golden suite
- 04 StepPublish SKILL.md
- 05 StepCanary the route
Must
- Trigger clear
- Golden pass or hold
- Must-not includes 'do not invent completion'
Must not
- Re-explain the style in chat
- Ship without a prompt version
# Generative AI Engineer id: generative-ai-engineer layer: app (App) second family: ai-trainer-evaluator tools: #/schema ## Pain The hired Generative AI Engineer seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - schema-pass · Generative AI Engineer analog close · analog #/schema · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
incident · bar 88
AI Security Specialist · MLOps Engineer
Triage, root-cause, patch, certify, hold canary until independent merge
feature · bar 85
Prompt Engineer · AI Trainer and Evaluator · AI Product Manager
Spec → implement → prove → independent certify → merge
chore · bar 82
MLOps Engineer · Data Engineer for AI
Low-risk maintenance under standard gates
pr · bar 88
AI Trainer and Evaluator · AI Security Specialist
Review PR, run eval gate, independent fleet, no self-approve
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 78. C1.
- Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
- Deny. credentials.reveal · deploy.canary
- Scopes. workspace_read
- Swarm seats. incident · feature · chore · pr
Daily missions
Run against live engines. Not slides.
skill.discover · skill.from_task
Package a generation workflow as Skill.md
Turn a repetitive prompt pack into a durable skill
Accept: SKILL.md published · Trigger clear. Pain removed: Stop re-explaining the same style guide. Surfaces: Debug Harness · Systems Lab.
Superpowers
- STE output standard
- Media ABR lab
- Skill.md packaging
Daily jobs
- Version a prompt
- Run golden suite
- Canary a model route
- Track token spend
Skills
- Prompting
- Fine-tuning
- Multimodal
- Diffusion
Toolkit
Command Center becomes this desk.
Command Center
Version prompt
Register system prompt + golden score
Outcome: Scored prompt version · caps.prompt.
Eval Gate
Run golden suite
Offline regression before ship
Outcome: Pass/fail board · caps.golden.
Models & Keys
Route model
Cheap vs reason cascade
Outcome: Model pick · caps.route.
Work templates
pr
GenAI feature PR
Ship generative UI feature with golden eval and canary
Paste: Feature brief + sample prompts. Accept: Prompt versioned · Golden pass · Canary decision.
Surfaces
Nav this desk actually opens.
Local analog
Lexical latent explorer
Hash-ngram anchors. Lexical manifold, not CLIP.
Lexical latent explorer (hash-ngram, not CLIP) and schema harden (JSON.parse required keys). SAMPLE.
Replacement cost
$165,000–$230,000 cash
Year-1 loaded + recruiting $292,300. Same band as LLM. Diffusion / latent desks.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $292,300 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.
Generative neighbors
Generative specialist
Prompt Engineer
System prompt that lives in Slack
A system prompt that lives in Slack. Two variants, no goldens, a judge that is the author.
Registry, STE, A/B through the gate. Chat history is not a prompt product.
- Tighten a flaky system prompt against goldens
- STE-clean a model reply for a production UI
- Run a regression suite on multi-turn tasks
Generative specialist
RAG Engineer
Hallucinated SKU in the FAQ — no chunk named
A RAG claim with no split, no overlap, no token count. Hallucinated product answers in the FAQ.
Named split, overlap, token count, 11-node graph. A vector demo is not retrieval.
- Fix a hallucinated product answer with a confidence gate
- Add metadata filters for a tenant
- Compare hybrid vs vectorless on hard queries
Generative specialist
NLP Engineer
Silent truncation on a 32k chat
Silent truncation on a long chat. Extraction F1 that nobody can reproduce.
Context pack under a named token budget. The triad scores the extract.
- Score the triad on hard queries
- Version extraction prompts against goldens
- Compress long chat under a named token budget
Generative specialist
Computer Vision Engineer
Detector that never loaded weights
A detector that never loaded weights. A CLIP sticker on a page that never ran a frame.
Sobel analog here. ViT/CLIP attachable. Goldens and drift before canary. A CLIP sticker is not vision ops.
- Check drift on a labeled set
- Canary weights behind a quality floor
- Pick edge topology when the plant is offline