The pain — unique to this desk
Foundation specialist · Research
LLM Engineer
Fluent answer with no citation and no fallback
What this agent actually does
RAG + routing under gates.
A fluent answer with no citation and no fallback. One frontier call is the whole plan.
The need this desk closes
Cited claims, a pattern scorecard, low-confidence flags, and a cheap-then-reason chain.
Why a generic chat cannot fake this
Fluent answer with no citation and no fallback
Cited claims and a cheap-then-reason chain. Fluency is not a product.
Takes
- Question
- Doc snippets or corpus notes
Returns
- Cited answer
- Pattern scorecard
- Low-confidence flags
Modalities this desk actually handles
How this specialist fuses
Fluent answer with no citation and no fallback
Question encoding + chunk encoding. Fusion is retrieval (hybrid / guarded / vectorless), then the LLM. Groundedness is a triad, not a vibe.
Vision-language on this desk
Named family. Never assumed.
If the question is about an image, Computer Vision encodes first (Sobel analog here; attached VLM — CLIP retrieve or LLaVA/Flamingo caption). This desk still requires citations. A caption without a chunk is a hold.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
A customer question hits the desk. rag.grounded returns claims plus citations. rag.pattern runs hybrid / guarded / vectorless and writes a pattern scorecard. Low-confidence claims flag instead of shipping fluency. Context hydration is a heatmap analog on a 2k–2M window — needle offset, not a model measurement. Fallback is a chain (cheaper model → cache → hold), not a single frontier call.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as Research. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- LLM Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #context. Kernel #/context on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
Answer this customer question from the product FAQ with citations. Flag any claim we cannot support. Do not call the frontier if the cheap route is faithful.
Sit keys llm / context / window / KV / LoRA / serving. Ticket kinds support and feature. RAG Engineer owns chunk strategy; this desk owns the serving loop and the fallback chain.
Playbook
How this agent works the ticket.
- 01 Groundrag.grounded: claims + citations on the pasted corpus.
- 02 Patternhybrid / guarded / vectorless. Pattern log shows the fusion path.
- 03 FlagUnsupported sentences hold. Fluency is not evidence.
- 04 Budget the windowContext pack under token budget. Silent truncation is a bug.
- 05 RouteCheap-then-reason cascade. Cost-aware, not all-frontier.
- 06 Eval11-method gate. Merge / hold / reject with reasons.
Acts like the role
Work it does. Work it will not fake.
Does
- Ship a grounded FAQ answer a customer can be handed
- Tune the pattern on a failing query
- Add a RAG pattern to the runbook with a scorecard
- Keep a fallback chain so one model outage is not an outage
Does not
- Call a hosted embedder from the sales page
- Treat the context heatmap as a model measurement
- Invent vLLM capacity on this host
- Grade its own answers — Trainer / Evaluator is the second family
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P0
RAG triad scoring
Today. Fluent answers without faithfulness checks
This agent. Faithfulness + context relevance + answer relevance
P0
Claim-level citations
Today. Only 50% of sentences supported in audits
This agent. Per-claim support flags + hold if unsupported
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P1
Local + frontier dual path
Today. All-or-nothing cloud dependency
This agent. Model vault with local fallback route
P1
Episodic → lessons consolidation
Today. Same failures every week
This agent. CoALA memory + procedural skill publish
P0
Typed MCP tool boundary
Today. N×M custom tool integrations
This agent. MCP registry + deny-list + audit per tool call
Skill.md
Customer corpus Q&A
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(llm-engineer).
Trigger. A question must be answered from product docs with citations.
- 01 StepTake the question and the doc snippets
- 02 StepRun grounded RAG
- 03 StepFlag low-confidence claims
- 04 StepScore through the eval gate
- 05 StepOnly then hand the answer to support
Must
- Citations present
- Low-confidence claims flagged
- Eval score recorded
Must not
- Answer from parametric memory when the corpus is silent
- Skip the fallback chain
# LLM Engineer id: llm-engineer layer: research (Research) second family: ai-trainer-evaluator tools: #/context ## Pain The hired LLM Engineer seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - context-pass · LLM Engineer analog close · analog #/context · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
Not named on a ticket-kind composition. Talk Route can still sit it from sit keys.
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 78. C1.
- Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
- Deny. credentials.reveal · deploy.canary
- Scopes. workspace_read
- Swarm seats. Not named on a ticket-kind composition. Talk Route can still sit it.
Daily missions
Run against live engines. Not slides.
rag.grounded · rag.pattern
Ship citation-aware RAG answer
Grounded answer with claim support + citations
Accept: Citations present · Low-confidence claims flagged. Pain removed: Hallucination reviews become mechanical. Surfaces: Intelligence Lab · Systems Lab.
Superpowers
- 15 RAG patterns
- Fallback model chain
- Context engineering pack
Daily jobs
- Ship a grounded FAQ answer
- Tune chunking on a failing query
- Add a new RAG pattern to the runbook
Skills
- Embeddings
- RAG
- Eval
- Prompt systems
Toolkit
Command Center becomes this desk.
Intelligence Lab
Grounded answer
Answer with claims + citations on your corpus
Outcome: Answer you can hand to a customer · rag.grounded.
Systems Lab
RAG pattern run
Pick hybrid / guarded / vectorless
Outcome: Pattern scorecard · rag.pattern.
Work templates
ticket
Customer corpus Q&A
Answer customer question from product docs with citations
Paste: Question + doc snippets or corpus notes. Accept: Citations · Low-confidence claims flagged · Eval score.
Surfaces
Nav this desk actually opens.
Local analog
Context hydration heatmap
Needle offset on a 2k–2M window. Heuristic, not a model measurement.
Context hydration heatmap. Needle offset on a 2k–2M window. Heuristic, not a model measurement.
Replacement cost
$165,000–$230,000 cash
Year-1 loaded + recruiting $292,300. LLM/RAG premium vs classical ML.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $292,300 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.
Foundation neighbors
Foundation specialist
AI Research Scientist
Paper claim that never left the notebook
A paper claim that never leaves the notebook. A diagram with a missing dim treated as compiled.
A scored experiment card with env classification. A chatbot summary is not a prototype.
- Reproduce a paper claim with a small prototype and an eval score
- Compare two retrieval heads on a pasted corpus
- Classify harness-layer failures without a model swap
Foundation specialist
ML Engineer
Notebook that never ships a scorer
A glue notebook that never ships a scorer. A green dashboard on two series nobody compared.
Confusion matrix plus PSI/KS before canary. A notebook pickle is not a model.
- Retrain on last week's labeled tickets and print confusion
- Canary a scorer behind the eval gate
- Profile features and flag drift on pasted series