Foundation specialist · Scale

ML Engineer

Notebook that never ships a scorer

What this agent actually does

Train → canary → watch.

A glue notebook that never ships a scorer. A green dashboard on two series nobody compared.

The pain — unique to this desk

A glue notebook that never ships a scorer. A green dashboard on two series nobody compared.

The need this desk closes

Accuracy plus confusion on real labels. PSI/KS named. Canary only after the eval family merges.

Why a generic chat cannot fake this

Notebook that never ships a scorer

Confusion matrix plus PSI/KS before canary. A notebook pickle is not a model.

Takes

  • Incident texts + labels
  • Two feature series

Returns

  • Accuracy + confusion
  • PSI / KS analog
  • Deploy decision

Modalities this desk actually handles

Labeled textNumeric seriesTabular features

How this specialist fuses

Notebook that never ships a scorer

Series and labels stay separate encodings. Drift is PSI / KS on the numbers; the classifier is fit on the labels. No image encoder unless CV sits.

Vision-language on this desk

Named family. Never assumed.

Vision-language is not this desk. If labels come from a VLM captioner, Trainer / Evaluator goldens the captions first.

Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.

Last week's labeled tickets arrive. ml.train fits logistic / RF / stump models in the in-browser classifier lab and reports accuracy plus confusion. Analytics profiles features. PSI and KS analogs compare two numeric series — they are not a production monitor. k8s.deploy is blocked unless the eval family merges. Drift is a job, not a green dashboard.

Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.

Reports up as Scale. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.

Five analog ticket kinds

How a hired human spends the week. What the analog kernel closes.

Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.

  1. ML Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #psi. Kernel #/psi on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

A typical ticket

What you say. What sits.

Train an incident labeler on these 200 tagged tickets. Hold the canary if confusion on SEV-1 is worse than last week.

Sit keys train / loss / dataset / feature / drift / PSI / KS. Ticket kinds chore and incident. MLOps owns the canary path. This desk trains; it does not courtesy-pass its own scorer.

Playbook

How this agent works the ticket.

  1. 01 Take labelsIncident-labeler template needs sample texts + labels. No labels, no train.
  2. 02 Fitml.train on the classifier lab. Accuracy and confusion are the artifact.
  3. 03 ProfileOLAP-style feature profiling in analytics. Mean, variance, Jensen–Shannon on two series.
  4. 04 Drift samplePSI / KS analogs on bins you paste. Not a live probe.
  5. 05 GateEval merge or hold. Authors never grade the model they just fit.
  6. 06 Canaryk8s.deploy only after the second family merges. Probe trail required.

Acts like the role

Work it does. Work it will not fake.

Does

  • Retrain on last week's labeled tickets and print confusion
  • Canary a scorer behind the eval gate
  • Profile features and flag drift on pasted series
  • Keep silent regressions off the path to prod

Does not

  • Invent a hosted MLflow or feature store
  • Paint two pasted bins as a production monitor
  • Bypass eval to ship a notebook pickle
  • Reveal credentials

Jobs on the board

Scenarios this desk was built to close.

From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.

P0

Golden dataset regression

Today. Prompt/model change ships without suite

This agent. Versioned golden set + offline gate before merge

P1

Production drift detect

Today. Model gets quietly worse

This agent. Drift monitor on labels/scores + retrain trigger

P0

Canary with quality SLO

Today. Infra canary ignores answer quality

This agent. Shadow traffic score vs baseline before promote

P1

Cost-aware model routing

Today. Single frontier model burns budget

This agent. Token budget + cascade cheap→frontier

P1

Corpus / feature freshness

Today. Stale RAG chunks and features

This agent. Ingest pipeline status + reindex job

P0

False completion graders

Today. Agent claims done while tests fail

This agent. Objective graders before done

P0

Prompt git-style versioning

Today. Prompts live in Notion / chat history

This agent. Registry with A/B + golden score

P1

Local + frontier dual path

Today. All-or-nothing cloud dependency

This agent. Model vault with local fallback route

Skill.md

Incident labeler

Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ml-engineer).

Trigger. A batch of labeled ops notes needs a gated scorer.

  1. 01 StepConfirm labels exist
  2. 02 StepFit the classifier lab
  3. 03 StepRecord accuracy and confusion
  4. 04 StepCompare PSI/KS against last week's bins
  5. 05 StepAsk eval to merge or hold before any canary

Must

  • Confusion known
  • Eval decision recorded

Must not

  • Ship on train-set accuracy
  • Treat PSI analog as a live monitor
# ML Engineer

id: ml-engineer
layer: scale (Scale)
second family: ai-trainer-evaluator
tools: #/psi

## Pain
The hired ML Engineer seat does not exist yet, or the work has no named ticket.

## Need
A named ticket, an analog close, and a second family. Not a chat that grades itself.

## Isolation
Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk.

## spend.halt
spend.halt on the ticket cap. Only the operator raises the ceiling.

## Ticket kinds
- psi-pass · ML Engineer analog close · analog #/psi · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

Who sits with this desk

Swarm compose by ticket kind.

Not named on a ticket-kind composition. Talk Route can still sit it from sit keys.

Sit keys

Talk Route matches these words.

psiksdrift

Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.

RBAC

Pass bar 78. C1.

  • Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
  • Deny. credentials.reveal · deploy.canary
  • Scopes. workspace_read
  • Swarm seats. Not named on a ticket-kind composition. Talk Route can still sit it.

Daily missions

Run against live engines. Not slides.

ml.train

Train ops incident classifier

Fit logistic / RF / XGBoost-style stumps on synthetic ops data

Accept: Accuracy reported · Confusion known. Pain removed: Instant train loop without notebook glue. Surfaces: Intelligence Lab.

Superpowers

  • In-browser classifier lab (7 algos)
  • OLAP feature profiling
  • Canary deploy gates

Daily jobs

  • Retrain on last week’s labeled tickets
  • Canary a scorer behind the eval gate
  • Profile features in analytics

Skills

  • ML
  • Statistics
  • Feature eng
  • Serving

Toolkit

Command Center becomes this desk.

Intelligence Lab

Train classifier lab

Fit models on your labeled ops/product data notes

Outcome: Accuracy + confusion for a real label set · ml.train.

Kubernetes

Canary deploy

Ship only after eval merge

Outcome: Release with probe trail · k8s.deploy.

Work templates

eval

Incident labeler

Train and gate a model that labels production incidents

Paste: Sample incident texts + labels. Accept: Model score · Eval merge or hold · Deploy decision.

Surfaces

Nav this desk actually opens.

Command CenterModels & KeysRole StudioIntelligence LabEvent AnalyticsKubernetesEval GateCertify & OperateUpdates

Local analog

Feature drift sandbox

Two numeric series. Mean, variance, Jensen–Shannon. Not a production monitor.

Feature drift sandbox, PSI, and KS. Two series. Not a production monitor. SAMPLE.

Run Feature drift sandbox

Replacement cost

$149,000$219,000 cash

Year-1 loaded + recruiting $272,320. KORE1 mid. Senior cash $220k–$300k+.. Not EngOS payroll. Not ARR.

Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $272,320 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.

Open Talk RouteAll twenty desks

AI Research ScientistLLM Engineer

EngOS