Generative specialist · App

RAG Engineer

Hallucinated SKU in the FAQ — no chunk named

What this agent actually does

Production retrieval.

A RAG claim with no split, no overlap, no token count. Hallucinated product answers in the FAQ.

The pain — unique to this desk

A RAG claim with no split, no overlap, no token count. Hallucinated product answers in the FAQ.

The need this desk closes

Named chunk strategy, logged fusion path, confidence gate. Hallucinations fail the gate, not the chat.

Why a generic chat cannot fake this

Hallucinated SKU in the FAQ — no chunk named

Named split, overlap, token count, 11-node graph. A vector demo is not retrieval.

Takes

  • KB excerpts
  • User question

Returns

  • Gated answer with citations
  • Fusion path log
  • Chunk audit

Modalities this desk actually handles

DocumentsText

How this specialist fuses

Hallucinated SKU in the FAQ — no chunk named

Document → chunk encoder (sliding / paragraph / recursive). Query encoder is the same analog. Fusion is hybrid + rerank, logged. Overlap % and token estimate are named.

Vision-language on this desk

Named family. Never assumed.

Multimodal RAG (image pages, screenshots) sits CV to encode the page — ViT patches if a VLM is attached — then this desk chunks the extracted text. We do not embed pixels into a vector DB on this host.

Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.

A hallucinated product answer is the ticket. rag.pipeline11 is the production graph with jailbreak plus confidence. rag.vectorless walks structure first when embeddings would lie. rag.pattern logs fusion. Chunk playground is the local analog: sliding / paragraph / recursive, overlap %, token estimate. Hallucinated product answers fail the gate, not the chat. Metadata filters for a tenant are a daily job.

Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.

Reports up as App. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.

Five analog ticket kinds

How a hired human spends the week. What the analog kernel closes.

Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.

  1. RAG Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #chunk. Kernel #/chunk on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

A typical ticket

What you say. What sits.

This FAQ answer cited a SKU we do not sell. Re-chunk, run hybrid vs vectorless, and hold if confidence is below the bar.

Sit keys rag / retrieve / chunk / embed / corpus. Ticket kinds support and research. LLM Engineer owns serving; Data Engineer owns freshness; this desk owns the retrieval stack.

Playbook

How this agent works the ticket.

  1. 01 Name the splitChunk playground: sliding / paragraph / recursive. Overlap % and token estimate required.
  2. 02 11-node graphrag.pipeline11: jailbreak node + confidence gate on the answer.
  3. 03 VectorlessStructure-first path through docs when the embedder would retrieve the wrong neighbor.
  4. 04 TriadFaithfulness + context relevance + answer relevance. Fluency is not a score.
  5. 05 Cite claimsPer-claim support flags. Unsupported sentences hold.
  6. 06 Cache-asideDual-store analog. Cache hit measured, not assumed.

Acts like the role

Work it does. Work it will not fake.

Does

  • Fix a hallucinated product answer with a confidence gate
  • Add metadata filters for a tenant
  • Compare hybrid vs vectorless on hard queries
  • Put split, overlap, and token count on the claim

Does not

  • Invent Pinecone / Weaviate / FAISS as live on this host
  • Ship a RAG claim with no split
  • Treat the chunk playground as a hosted embedder
  • Let a jailbreak in retrieved docs call a tool

Jobs on the board

Scenarios this desk was built to close.

From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.

P0

RAG triad scoring

Today. Fluent answers without faithfulness checks

This agent. Faithfulness + context relevance + answer relevance

P0

Claim-level citations

Today. Only 50% of sentences supported in audits

This agent. Per-claim support flags + hold if unsupported

P1

Corpus / feature freshness

Today. Stale RAG chunks and features

This agent. Ingest pipeline status + reindex job

P0

Golden dataset regression

Today. Prompt/model change ships without suite

This agent. Versioned golden set + offline gate before merge

P0

Direct + indirect prompt injection

Today. OWASP #1 — injection via user and retrieved docs

This agent. Red-team pack + input/tool/output layers

P1

Cost-aware model routing

Today. Single frontier model burns budget

This agent. Token budget + cascade cheap→frontier

P0

False completion graders

Today. Agent claims done while tests fail

This agent. Objective graders before done

P0

Canary with quality SLO

Today. Infra canary ignores answer quality

This agent. Shadow traffic score vs baseline before promote

Skill.md

KB answer SLA

Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(rag-engineer).

Trigger. A knowledge-base question must meet a groundedness threshold.

  1. 01 StepTake KB excerpts + the user question
  2. 02 StepChoose chunk strategy and name overlap
  3. 03 StepRun the 11-node graph
  4. 04 StepRequire citations and a confidence gate
  5. 05 StepBlock jailbreak in retrieved text

Must

  • Citations
  • Confidence gate
  • No jailbreak

Must not

  • Answer without a split
  • Skip vectorless when hybrid retrieves the wrong neighbor
# RAG Engineer

id: rag-engineer
layer: app (App)
second family: ai-trainer-evaluator
tools: #/chunk

## Pain
The hired RAG Engineer seat does not exist yet, or the work has no named ticket.

## Need
A named ticket, an analog close, and a second family. Not a chat that grades itself.

## Isolation
Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk.

## spend.halt
spend.halt on the ticket cap. Only the operator raises the ceiling.

## Ticket kinds
- chunk-pass · RAG Engineer analog close · analog #/chunk · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

Who sits with this desk

Swarm compose by ticket kind.

Not named on a ticket-kind composition. Talk Route can still sit it from sit keys.

Sit keys

Talk Route matches these words.

chunk

Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.

RBAC

Pass bar 78. C1.

  • Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
  • Deny. credentials.reveal · deploy.canary
  • Scopes. workspace_read
  • Swarm seats. Not named on a ticket-kind composition. Talk Route can still sit it.

Daily missions

Run against live engines. Not slides.

rag.pattern · rag.query_advanced

Route multi-index retrieval

Run routing + hybrid + rerank pattern chain

Accept: Pattern log shows fusion path · Cache hit measured. Pain removed: One UI for the whole retrieval stack. Surfaces: Systems Lab · Intelligence Lab.

Superpowers

  • 15 production RAG patterns
  • Graph k-hop
  • Dual-store cache-aside

Daily jobs

  • Fix a hallucinated product answer
  • Add metadata filters for a tenant
  • Compare hybrid vs vectorless on hard queries

Skills

  • Embeddings
  • Search
  • Retrieval
  • Chunking

Toolkit

Command Center becomes this desk.

Agent Factory

11-node RAG graph

Full production pipeline with jailbreak + confidence

Outcome: Gated answer with confidence · rag.pipeline11.

Intelligence Lab

Vectorless navigate

Structure-first retrieval

Outcome: Path through docs without embeddings · rag.vectorless.

Work templates

ticket

KB answer SLA

Answer from customer knowledge base with groundedness > threshold

Paste: KB excerpts + user question. Accept: Citations · Confidence gate · No jailbreak.

Surfaces

Nav this desk actually opens.

Command CenterModels & KeysRole StudioIntelligence LabAgent FactorySystems LabEval GateAgent MemoryCertify & OperateUpdates

Local analog

Chunk strategy playground

Sliding, paragraph, or recursive split. Overlap %, token estimate.

Chunk strategy playground. Sliding, paragraph, or recursive. Overlap %, token estimate. SAMPLE.

Run Chunk strategy playground

Replacement cost

$155,000$220,000 cash

Year-1 loaded + recruiting $277,500. NLP/LLM adjacent. Not a published BLS SOC.. Not EngOS payroll. Not ARR.

Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $277,500 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.

Open Talk RouteAll twenty desks

Prompt EngineerAI Agent Engineer

Generative neighbors

Generative specialist

Generative AI Engineer

Playground demo sold as a generation system

A playground demo sold as a generation system. The style guide lives in chat history.

Skill.md, goldens, and a cost route. A pretty sample is not a system.

TextImage (lexical latent analog)Video (ABR lab)
  • Version a prompt and keep the score with it
  • Run a golden suite before a generative UI ships
  • Package a repetitive style guide as Skill.md
Analog · Lexical latent explorer · bar 78 · C1

Generative specialist

Prompt Engineer

System prompt that lives in Slack

A system prompt that lives in Slack. Two variants, no goldens, a judge that is the author.

Registry, STE, A/B through the gate. Chat history is not a prompt product.

Text
  • Tighten a flaky system prompt against goldens
  • STE-clean a model reply for a production UI
  • Run a regression suite on multi-turn tasks
Analog · Tokenizer & priority · bar 78 · C1

Generative specialist

NLP Engineer

Silent truncation on a 32k chat

Silent truncation on a long chat. Extraction F1 that nobody can reproduce.

Context pack under a named token budget. The triad scores the extract.

Text
  • Score the triad on hard queries
  • Version extraction prompts against goldens
  • Compress long chat under a named token budget
Analog · Embedding distance · bar 78 · C1

Generative specialist

Computer Vision Engineer

Detector that never loaded weights

A detector that never loaded weights. A CLIP sticker on a page that never ran a frame.

Sobel analog here. ViT/CLIP attachable. Goldens and drift before canary. A CLIP sticker is not vision ops.

ImageVideo (ABR ladder)Edge topology
  • Check drift on a labeled set
  • Canary weights behind a quality floor
  • Pick edge topology when the plant is offline
Analog · Edge overlay · bar 78 · C1

EngOS