Foundation specialist · Research

LLM Engineer

Fluent answer with no citation and no fallback

What this agent actually does

RAG + routing under gates.

A fluent answer with no citation and no fallback. One frontier call is the whole plan.

The pain — unique to this desk

A fluent answer with no citation and no fallback. One frontier call is the whole plan.

The need this desk closes

Cited claims, a pattern scorecard, low-confidence flags, and a cheap-then-reason chain.

Why a generic chat cannot fake this

Fluent answer with no citation and no fallback

Cited claims and a cheap-then-reason chain. Fluency is not a product.

Takes

  • Question
  • Doc snippets or corpus notes

Returns

  • Cited answer
  • Pattern scorecard
  • Low-confidence flags

Modalities this desk actually handles

TextCorpus documents

How this specialist fuses

Fluent answer with no citation and no fallback

Question encoding + chunk encoding. Fusion is retrieval (hybrid / guarded / vectorless), then the LLM. Groundedness is a triad, not a vibe.

Vision-language on this desk

Named family. Never assumed.

If the question is about an image, Computer Vision encodes first (Sobel analog here; attached VLM — CLIP retrieve or LLaVA/Flamingo caption). This desk still requires citations. A caption without a chunk is a hold.

Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.

A customer question hits the desk. rag.grounded returns claims plus citations. rag.pattern runs hybrid / guarded / vectorless and writes a pattern scorecard. Low-confidence claims flag instead of shipping fluency. Context hydration is a heatmap analog on a 2k–2M window — needle offset, not a model measurement. Fallback is a chain (cheaper model → cache → hold), not a single frontier call.

Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.

Reports up as Research. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.

Five analog ticket kinds

How a hired human spends the week. What the analog kernel closes.

Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.

  1. LLM Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #context. Kernel #/context on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

A typical ticket

What you say. What sits.

Answer this customer question from the product FAQ with citations. Flag any claim we cannot support. Do not call the frontier if the cheap route is faithful.

Sit keys llm / context / window / KV / LoRA / serving. Ticket kinds support and feature. RAG Engineer owns chunk strategy; this desk owns the serving loop and the fallback chain.

Playbook

How this agent works the ticket.

  1. 01 Groundrag.grounded: claims + citations on the pasted corpus.
  2. 02 Patternhybrid / guarded / vectorless. Pattern log shows the fusion path.
  3. 03 FlagUnsupported sentences hold. Fluency is not evidence.
  4. 04 Budget the windowContext pack under token budget. Silent truncation is a bug.
  5. 05 RouteCheap-then-reason cascade. Cost-aware, not all-frontier.
  6. 06 Eval11-method gate. Merge / hold / reject with reasons.

Acts like the role

Work it does. Work it will not fake.

Does

  • Ship a grounded FAQ answer a customer can be handed
  • Tune the pattern on a failing query
  • Add a RAG pattern to the runbook with a scorecard
  • Keep a fallback chain so one model outage is not an outage

Does not

  • Call a hosted embedder from the sales page
  • Treat the context heatmap as a model measurement
  • Invent vLLM capacity on this host
  • Grade its own answers — Trainer / Evaluator is the second family

Jobs on the board

Scenarios this desk was built to close.

From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.

P0

RAG triad scoring

Today. Fluent answers without faithfulness checks

This agent. Faithfulness + context relevance + answer relevance

P0

Claim-level citations

Today. Only 50% of sentences supported in audits

This agent. Per-claim support flags + hold if unsupported

P0

Golden dataset regression

Today. Prompt/model change ships without suite

This agent. Versioned golden set + offline gate before merge

P1

Cost-aware model routing

Today. Single frontier model burns budget

This agent. Token budget + cascade cheap→frontier

P0

Prompt git-style versioning

Today. Prompts live in Notion / chat history

This agent. Registry with A/B + golden score

P1

Local + frontier dual path

Today. All-or-nothing cloud dependency

This agent. Model vault with local fallback route

P1

Episodic → lessons consolidation

Today. Same failures every week

This agent. CoALA memory + procedural skill publish

P0

Typed MCP tool boundary

Today. N×M custom tool integrations

This agent. MCP registry + deny-list + audit per tool call

Skill.md

Customer corpus Q&A

Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(llm-engineer).

Trigger. A question must be answered from product docs with citations.

  1. 01 StepTake the question and the doc snippets
  2. 02 StepRun grounded RAG
  3. 03 StepFlag low-confidence claims
  4. 04 StepScore through the eval gate
  5. 05 StepOnly then hand the answer to support

Must

  • Citations present
  • Low-confidence claims flagged
  • Eval score recorded

Must not

  • Answer from parametric memory when the corpus is silent
  • Skip the fallback chain
# LLM Engineer

id: llm-engineer
layer: research (Research)
second family: ai-trainer-evaluator
tools: #/context

## Pain
The hired LLM Engineer seat does not exist yet, or the work has no named ticket.

## Need
A named ticket, an analog close, and a second family. Not a chat that grades itself.

## Isolation
Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk.

## spend.halt
spend.halt on the ticket cap. Only the operator raises the ceiling.

## Ticket kinds
- context-pass · LLM Engineer analog close · analog #/context · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

Who sits with this desk

Swarm compose by ticket kind.

Not named on a ticket-kind composition. Talk Route can still sit it from sit keys.

Sit keys

Talk Route matches these words.

context

Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.

RBAC

Pass bar 78. C1.

  • Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
  • Deny. credentials.reveal · deploy.canary
  • Scopes. workspace_read
  • Swarm seats. Not named on a ticket-kind composition. Talk Route can still sit it.

Daily missions

Run against live engines. Not slides.

rag.grounded · rag.pattern

Ship citation-aware RAG answer

Grounded answer with claim support + citations

Accept: Citations present · Low-confidence claims flagged. Pain removed: Hallucination reviews become mechanical. Surfaces: Intelligence Lab · Systems Lab.

Superpowers

  • 15 RAG patterns
  • Fallback model chain
  • Context engineering pack

Daily jobs

  • Ship a grounded FAQ answer
  • Tune chunking on a failing query
  • Add a new RAG pattern to the runbook

Skills

  • Embeddings
  • RAG
  • Eval
  • Prompt systems

Toolkit

Command Center becomes this desk.

Intelligence Lab

Grounded answer

Answer with claims + citations on your corpus

Outcome: Answer you can hand to a customer · rag.grounded.

Systems Lab

RAG pattern run

Pick hybrid / guarded / vectorless

Outcome: Pattern scorecard · rag.pattern.

Work templates

ticket

Customer corpus Q&A

Answer customer question from product docs with citations

Paste: Question + doc snippets or corpus notes. Accept: Citations · Low-confidence claims flagged · Eval score.

Surfaces

Nav this desk actually opens.

Command CenterModels & KeysRole StudioIntelligence LabSystems LabEval GateAgent MemoryAgent FactoryCertify & OperateUpdates

Local analog

Context hydration heatmap

Needle offset on a 2k–2M window. Heuristic, not a model measurement.

Context hydration heatmap. Needle offset on a 2k–2M window. Heuristic, not a model measurement.

Run Context hydration heatmap

Replacement cost

$165,000$230,000 cash

Year-1 loaded + recruiting $292,300. LLM/RAG premium vs classical ML.. Not EngOS payroll. Not ARR.

Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $292,300 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.

Open Talk RouteAll twenty desks

ML EngineerGenerative AI Engineer

EngOS