The pain — unique to this desk
Business specialist · App
AI Product Manager
Feature brief with no eval plan and no halt
What this agent actually does
Ship AI features with evidence.
A feature brief with no eval plan and no halt. A demo GIF treated as a success metric.
The need this desk closes
Brief, named success metric, eval plan, published cap, spend.halt. Token math on numbers you supply.
Why a generic chat cannot fake this
Feature brief with no eval plan and no halt
Named success metric, eval plan, published cap, spend.halt. A demo GIF is not a metric.
Takes
- User problem + constraints + success metric
Returns
- Brief artifact
- Eval plan
- Economics analog
Modalities this desk actually handles
How this specialist fuses
Feature brief with no eval plan and no halt
Success metric + eval plan + halt. A multimodal feature names the encoder and the judge in the brief.
Vision-language on this desk
Named family. Never assumed.
If the feature is vision-language, the brief requires image goldens and a second family. A demo GIF is not an eval plan.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
A one-liner arrives. Command Center incubation is research → plan → review. The brief artifact has success metrics, an eval plan, and risk notes. findings.fix_all is the launch checklist. Token unit economics multiplies the $/1M and HITL cost you supply — no vendor price list invented here. Published caps. spend.halt on over-cap. Role capacity and cert coverage are reviewed, not assumed.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as App. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- AI Product Manager analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #economics. Kernel #/economics on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
Scope citation chips for the FAQ bot. Success = groundedness ≥ bar on the golden set. Daily cap $50. Halt over cap. Launch checklist before GA.
Sit keys pricing / cap / seat / economics / roadmap. Ticket kinds feature, billing, support. Engineer builds. Eval gates. This desk accepts.
Playbook
How this agent works the ticket.
- 01 IncubateOne-liner through research → plan → review. Shared language with eng.
- 02 BriefUser problem, constraints, success metric, eval plan, risk notes.
- 03 EconomicsYou supply $/1M and HITL. We multiply. No invented vendor list.
- 04 CapPublished daily cap. spend.halt on over-cap.
- 05 Launch chainFix-all prod readiness. UX a11y chain if the surface is UI.
- 06 Cert coverageWhich desks are certified to touch this feature. Gaps are backlog.
Acts like the role
Work it does. Work it will not fake.
Does
- Write an AI feature brief with eval criteria
- Prioritize hardening findings
- Review agent cert coverage by role
- Run fix-all prod readiness as a launch checklist
- Publish a cap and a halt, not a hope
Does not
- Invent a vendor price list
- Ship AI without a measurable outcome
- Treat ARR as claimed (ARR_CLAIMED=false)
- Let the authoring desk grade the feature it wrote
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P1
Success metrics on AI features
Today. Ship AI without measurable outcomes
This agent. Feature brief + eval plan + KPI log
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P0
Canary with quality SLO
Today. Infra canary ignores answer quality
This agent. Shadow traffic score vs baseline before promote
P1
Decision audit trail
Today. Cannot explain model decisions to compliance
This agent. Trace + policy tags on every ship decision
P0
False completion graders
Today. Agent claims done while tests fail
This agent. Objective graders before done
P1
Incident runbook agent
Today. Pager chaos without structured triage
This agent. Incident desk with severity + rollback path
Skill.md
Feature brief
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ai-product-manager).
Trigger. An AI feature is being scoped and a success metric can be named.
- 01 StepWrite the user problem + constraints
- 02 StepName the success metric
- 03 StepAttach an eval plan
- 04 StepSet the cap and the halt
- 05 StepFile risk notes
Must
- Brief artifact
- Eval plan
- Risk notes
Must not
- Ship without a metric
- Invent prices
- Claim ARR
# AI Product Manager id: ai-product-manager layer: app (App) second family: ai-trainer-evaluator tools: #/economics ## Pain The hired AI Product Manager seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - economics-pass · AI Product Manager analog close · analog #/economics · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
support · bar 80
Forward Deployed Engineer · Prompt Engineer
Deflect with grounded corpus, escalate billing-sensitive cases
feature · bar 85
Generative AI Engineer · Prompt Engineer · AI Trainer and Evaluator
Spec → implement → prove → independent certify → merge
billing · bar 82
Forward Deployed Engineer · MLOps Engineer
Cost caps, plan, usage ledger — admin only paths
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 78. C1.
- Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
- Deny. credentials.reveal · deploy.canary
- Scopes. workspace_read
- Swarm seats. support · feature · billing
Daily missions
Run against live engines. Not slides.
graph.map
Incubate a product brief
Route a one-liner through Master Architect incubation
Accept: Buildout log advanced. Pain removed: PM ↔ eng shared language in one surface. Surfaces: Command Center · Agent Fleet.
Superpowers
- Command center KPIs
- Mission XP per role
- Incubation loop
Daily jobs
- Write an AI feature brief with eval criteria
- Prioritize hardening findings
- Review agent cert coverage by role
- Run fix-all prod readiness for launch checklist
Skills
- Strategy
- Roadmaps
- Analytics
- Evals
Toolkit
Command Center becomes this desk.
Command Center
Incubate product brief
Research → plan → review trail
Outcome: Decision-ready incubation log.
Kernel Guard
Platform audit
Full-spectrum hardening + fix-all score
Outcome: Prioritized backlog · findings.fix_all.
Kernel Guard
UX a11y chain
Sequential UI/accessibility overhaul prompts
Outcome: UX debt map · findings.fix_all.
Work templates
design
Feature brief
Scope an AI feature with success metrics and eval plan
Paste: User problem + constraints + success metric. Accept: Brief artifact · Eval plan · Risk notes.
Surfaces
Nav this desk actually opens.
Local analog
Token unit economics
You supply $/1M and HITL cost. We multiply. No vendor price list.
Token unit economics. You supply $/1M and HITL cost. We multiply. No vendor price list. SAMPLE.
Replacement cost
$150,000–$220,000 cash
Year-1 loaded + recruiting $273,800. Senior PM plus AI premium. Equity excluded.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $273,800 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.