Business specialist · App

AI Product Manager

Feature brief with no eval plan and no halt

What this agent actually does

Ship AI features with evidence.

A feature brief with no eval plan and no halt. A demo GIF treated as a success metric.

The pain — unique to this desk

A feature brief with no eval plan and no halt. A demo GIF treated as a success metric.

The need this desk closes

Brief, named success metric, eval plan, published cap, spend.halt. Token math on numbers you supply.

Why a generic chat cannot fake this

Feature brief with no eval plan and no halt

Named success metric, eval plan, published cap, spend.halt. A demo GIF is not a metric.

Takes

  • User problem + constraints + success metric

Returns

  • Brief artifact
  • Eval plan
  • Economics analog

Modalities this desk actually handles

TextMetrics

How this specialist fuses

Feature brief with no eval plan and no halt

Success metric + eval plan + halt. A multimodal feature names the encoder and the judge in the brief.

Vision-language on this desk

Named family. Never assumed.

If the feature is vision-language, the brief requires image goldens and a second family. A demo GIF is not an eval plan.

Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.

A one-liner arrives. Command Center incubation is research → plan → review. The brief artifact has success metrics, an eval plan, and risk notes. findings.fix_all is the launch checklist. Token unit economics multiplies the $/1M and HITL cost you supply — no vendor price list invented here. Published caps. spend.halt on over-cap. Role capacity and cert coverage are reviewed, not assumed.

Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.

Reports up as App. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.

Five analog ticket kinds

How a hired human spends the week. What the analog kernel closes.

Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.

  1. AI Product Manager analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #economics. Kernel #/economics on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

A typical ticket

What you say. What sits.

Scope citation chips for the FAQ bot. Success = groundedness ≥ bar on the golden set. Daily cap $50. Halt over cap. Launch checklist before GA.

Sit keys pricing / cap / seat / economics / roadmap. Ticket kinds feature, billing, support. Engineer builds. Eval gates. This desk accepts.

Playbook

How this agent works the ticket.

  1. 01 IncubateOne-liner through research → plan → review. Shared language with eng.
  2. 02 BriefUser problem, constraints, success metric, eval plan, risk notes.
  3. 03 EconomicsYou supply $/1M and HITL. We multiply. No invented vendor list.
  4. 04 CapPublished daily cap. spend.halt on over-cap.
  5. 05 Launch chainFix-all prod readiness. UX a11y chain if the surface is UI.
  6. 06 Cert coverageWhich desks are certified to touch this feature. Gaps are backlog.

Acts like the role

Work it does. Work it will not fake.

Does

  • Write an AI feature brief with eval criteria
  • Prioritize hardening findings
  • Review agent cert coverage by role
  • Run fix-all prod readiness as a launch checklist
  • Publish a cap and a halt, not a hope

Does not

  • Invent a vendor price list
  • Ship AI without a measurable outcome
  • Treat ARR as claimed (ARR_CLAIMED=false)
  • Let the authoring desk grade the feature it wrote

Jobs on the board

Scenarios this desk was built to close.

From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.

P1

Success metrics on AI features

Today. Ship AI without measurable outcomes

This agent. Feature brief + eval plan + KPI log

P0

Golden dataset regression

Today. Prompt/model change ships without suite

This agent. Versioned golden set + offline gate before merge

P0

Prompt git-style versioning

Today. Prompts live in Notion / chat history

This agent. Registry with A/B + golden score

P1

Cost-aware model routing

Today. Single frontier model burns budget

This agent. Token budget + cascade cheap→frontier

P0

Canary with quality SLO

Today. Infra canary ignores answer quality

This agent. Shadow traffic score vs baseline before promote

P1

Decision audit trail

Today. Cannot explain model decisions to compliance

This agent. Trace + policy tags on every ship decision

P0

False completion graders

Today. Agent claims done while tests fail

This agent. Objective graders before done

P1

Incident runbook agent

Today. Pager chaos without structured triage

This agent. Incident desk with severity + rollback path

Skill.md

Feature brief

Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ai-product-manager).

Trigger. An AI feature is being scoped and a success metric can be named.

  1. 01 StepWrite the user problem + constraints
  2. 02 StepName the success metric
  3. 03 StepAttach an eval plan
  4. 04 StepSet the cap and the halt
  5. 05 StepFile risk notes

Must

  • Brief artifact
  • Eval plan
  • Risk notes

Must not

  • Ship without a metric
  • Invent prices
  • Claim ARR
# AI Product Manager

id: ai-product-manager
layer: app (App)
second family: ai-trainer-evaluator
tools: #/economics

## Pain
The hired AI Product Manager seat does not exist yet, or the work has no named ticket.

## Need
A named ticket, an analog close, and a second family. Not a chat that grades itself.

## Isolation
Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk.

## spend.halt
spend.halt on the ticket cap. Only the operator raises the ceiling.

## Ticket kinds
- economics-pass · AI Product Manager analog close · analog #/economics · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

Who sits with this desk

Swarm compose by ticket kind.

  • support · bar 80

    Forward Deployed Engineer · Prompt Engineer

    Deflect with grounded corpus, escalate billing-sensitive cases

  • feature · bar 85

    Generative AI Engineer · Prompt Engineer · AI Trainer and Evaluator

    Spec → implement → prove → independent certify → merge

  • billing · bar 82

    Forward Deployed Engineer · MLOps Engineer

    Cost caps, plan, usage ledger — admin only paths

Sit keys

Talk Route matches these words.

economics

Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.

RBAC

Pass bar 78. C1.

  • Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
  • Deny. credentials.reveal · deploy.canary
  • Scopes. workspace_read
  • Swarm seats. support · feature · billing

Daily missions

Run against live engines. Not slides.

graph.map

Incubate a product brief

Route a one-liner through Master Architect incubation

Accept: Buildout log advanced. Pain removed: PM ↔ eng shared language in one surface. Surfaces: Command Center · Agent Fleet.

Superpowers

  • Command center KPIs
  • Mission XP per role
  • Incubation loop

Daily jobs

  • Write an AI feature brief with eval criteria
  • Prioritize hardening findings
  • Review agent cert coverage by role
  • Run fix-all prod readiness for launch checklist

Skills

  • Strategy
  • Roadmaps
  • Analytics
  • Evals

Toolkit

Command Center becomes this desk.

Command Center

Incubate product brief

Research → plan → review trail

Outcome: Decision-ready incubation log.

Kernel Guard

Platform audit

Full-spectrum hardening + fix-all score

Outcome: Prioritized backlog · findings.fix_all.

Kernel Guard

UX a11y chain

Sequential UI/accessibility overhaul prompts

Outcome: UX debt map · findings.fix_all.

Work templates

design

Feature brief

Scope an AI feature with success metrics and eval plan

Paste: User problem + constraints + success metric. Accept: Brief artifact · Eval plan · Risk notes.

Surfaces

Nav this desk actually opens.

Command CenterModels & KeysRole StudioEval GateTalent & HireEvent AnalyticsAgent FleetKernel GuardCertify & OperateUpdates

Local analog

Token unit economics

You supply $/1M and HITL cost. We multiply. No vendor price list.

Token unit economics. You supply $/1M and HITL cost. We multiply. No vendor price list. SAMPLE.

Run Token unit economics

Replacement cost

$150,000$220,000 cash

Year-1 loaded + recruiting $273,800. Senior PM plus AI premium. Equity excluded.. Not EngOS payroll. Not ARR.

Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $273,800 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.

Open Talk RouteAll twenty desks

Data Engineer for AIAI Solutions Architect

EngOS