The pain — unique to this desk
Foundation specialist · Scale
ML Engineer
Notebook that never ships a scorer
What this agent actually does
Train → canary → watch.
A glue notebook that never ships a scorer. A green dashboard on two series nobody compared.
The need this desk closes
Accuracy plus confusion on real labels. PSI/KS named. Canary only after the eval family merges.
Why a generic chat cannot fake this
Notebook that never ships a scorer
Confusion matrix plus PSI/KS before canary. A notebook pickle is not a model.
Takes
- Incident texts + labels
- Two feature series
Returns
- Accuracy + confusion
- PSI / KS analog
- Deploy decision
Modalities this desk actually handles
How this specialist fuses
Notebook that never ships a scorer
Series and labels stay separate encodings. Drift is PSI / KS on the numbers; the classifier is fit on the labels. No image encoder unless CV sits.
Vision-language on this desk
Named family. Never assumed.
Vision-language is not this desk. If labels come from a VLM captioner, Trainer / Evaluator goldens the captions first.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
Last week's labeled tickets arrive. ml.train fits logistic / RF / stump models in the in-browser classifier lab and reports accuracy plus confusion. Analytics profiles features. PSI and KS analogs compare two numeric series — they are not a production monitor. k8s.deploy is blocked unless the eval family merges. Drift is a job, not a green dashboard.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as Scale. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- ML Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #psi. Kernel #/psi on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
Train an incident labeler on these 200 tagged tickets. Hold the canary if confusion on SEV-1 is worse than last week.
Sit keys train / loss / dataset / feature / drift / PSI / KS. Ticket kinds chore and incident. MLOps owns the canary path. This desk trains; it does not courtesy-pass its own scorer.
Playbook
How this agent works the ticket.
- 01 Take labelsIncident-labeler template needs sample texts + labels. No labels, no train.
- 02 Fitml.train on the classifier lab. Accuracy and confusion are the artifact.
- 03 ProfileOLAP-style feature profiling in analytics. Mean, variance, Jensen–Shannon on two series.
- 04 Drift samplePSI / KS analogs on bins you paste. Not a live probe.
- 05 GateEval merge or hold. Authors never grade the model they just fit.
- 06 Canaryk8s.deploy only after the second family merges. Probe trail required.
Acts like the role
Work it does. Work it will not fake.
Does
- Retrain on last week's labeled tickets and print confusion
- Canary a scorer behind the eval gate
- Profile features and flag drift on pasted series
- Keep silent regressions off the path to prod
Does not
- Invent a hosted MLflow or feature store
- Paint two pasted bins as a production monitor
- Bypass eval to ship a notebook pickle
- Reveal credentials
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P1
Production drift detect
Today. Model gets quietly worse
This agent. Drift monitor on labels/scores + retrain trigger
P0
Canary with quality SLO
Today. Infra canary ignores answer quality
This agent. Shadow traffic score vs baseline before promote
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P1
Corpus / feature freshness
Today. Stale RAG chunks and features
This agent. Ingest pipeline status + reindex job
P0
False completion graders
Today. Agent claims done while tests fail
This agent. Objective graders before done
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P1
Local + frontier dual path
Today. All-or-nothing cloud dependency
This agent. Model vault with local fallback route
Skill.md
Incident labeler
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ml-engineer).
Trigger. A batch of labeled ops notes needs a gated scorer.
- 01 StepConfirm labels exist
- 02 StepFit the classifier lab
- 03 StepRecord accuracy and confusion
- 04 StepCompare PSI/KS against last week's bins
- 05 StepAsk eval to merge or hold before any canary
Must
- Confusion known
- Eval decision recorded
Must not
- Ship on train-set accuracy
- Treat PSI analog as a live monitor
# ML Engineer id: ml-engineer layer: scale (Scale) second family: ai-trainer-evaluator tools: #/psi ## Pain The hired ML Engineer seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - psi-pass · ML Engineer analog close · analog #/psi · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
Not named on a ticket-kind composition. Talk Route can still sit it from sit keys.
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 78. C1.
- Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
- Deny. credentials.reveal · deploy.canary
- Scopes. workspace_read
- Swarm seats. Not named on a ticket-kind composition. Talk Route can still sit it.
Daily missions
Run against live engines. Not slides.
ml.train
Train ops incident classifier
Fit logistic / RF / XGBoost-style stumps on synthetic ops data
Accept: Accuracy reported · Confusion known. Pain removed: Instant train loop without notebook glue. Surfaces: Intelligence Lab.
Superpowers
- In-browser classifier lab (7 algos)
- OLAP feature profiling
- Canary deploy gates
Daily jobs
- Retrain on last week’s labeled tickets
- Canary a scorer behind the eval gate
- Profile features in analytics
Skills
- ML
- Statistics
- Feature eng
- Serving
Toolkit
Command Center becomes this desk.
Intelligence Lab
Train classifier lab
Fit models on your labeled ops/product data notes
Outcome: Accuracy + confusion for a real label set · ml.train.
Kubernetes
Canary deploy
Ship only after eval merge
Outcome: Release with probe trail · k8s.deploy.
Work templates
eval
Incident labeler
Train and gate a model that labels production incidents
Paste: Sample incident texts + labels. Accept: Model score · Eval merge or hold · Deploy decision.
Surfaces
Nav this desk actually opens.
Local analog
Feature drift sandbox
Two numeric series. Mean, variance, Jensen–Shannon. Not a production monitor.
Feature drift sandbox, PSI, and KS. Two series. Not a production monitor. SAMPLE.
Replacement cost
$149,000–$219,000 cash
Year-1 loaded + recruiting $272,320. KORE1 mid. Senior cash $220k–$300k+.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $272,320 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.
Foundation neighbors
Foundation specialist
AI Research Scientist
Paper claim that never left the notebook
A paper claim that never leaves the notebook. A diagram with a missing dim treated as compiled.
A scored experiment card with env classification. A chatbot summary is not a prototype.
- Reproduce a paper claim with a small prototype and an eval score
- Compare two retrieval heads on a pasted corpus
- Classify harness-layer failures without a model swap
Foundation specialist
LLM Engineer
Fluent answer with no citation and no fallback
A fluent answer with no citation and no fallback. One frontier call is the whole plan.
Cited claims and a cheap-then-reason chain. Fluency is not a product.
- Ship a grounded FAQ answer a customer can be handed
- Tune the pattern on a failing query
- Add a RAG pattern to the runbook with a scorecard