The pain — unique to this desk
Generative specialist · Research
NLP Engineer
Silent truncation on a 32k chat
What this agent actually does
Language systems with RAG rigor.
Silent truncation on a long chat. Extraction F1 that nobody can reproduce.
The need this desk closes
A context pack under a named token budget. Faithfulness, context relevance, answer relevance — the triad.
Why a generic chat cannot fake this
Silent truncation on a 32k chat
Context pack under a named token budget. The triad scores the extract.
Takes
- Samples + labels
- Long chat
Returns
- Faithfulness / context / answer scores
- Compressed context pack
Modalities this desk actually handles
How this specialist fuses
Silent truncation on a 32k chat
Token budget is the fuse: context.assemble layers sources until the window is honest. Subword tokenize analog is hash-ngram cosine — not a hosted embedder.
Vision-language on this desk
Named family. Never assumed.
OCR / caption text from a VLM arrives as text. This desk compresses and scores it. It does not run the vision encoder.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
A long chat is truncating silently. context.assemble packs under the token budget. caps.triad scores faithfulness, context relevance, answer relevance. Prompt registry versions the extraction prompt. Hash-ngram cosine, character n-grams, and embed distance are local analogs — not a hosted embedder. Corpus refresh is a job shared with Data Engineer.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as Research. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- NLP Engineer analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #embed. Kernel #/embed on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
This extraction F1 dropped on last week's tickets. Version the prompt, score the triad, and stop silent truncation on the 32k window.
Sit keys NER / classify / n-gram / nlp. Ticket kinds feature and pr. Prompt Engineer owns the system prompt; this desk owns language understanding and the context pack.
Playbook
How this agent works the ticket.
- 01 Budgetcontext.assemble. Budget respected. Sources layered. Silent truncation is a defect.
- 02 TriadFaithfulness board on hard queries.
- 03 VersionNLP prompts go in the registry with a golden delta.
- 04 DistanceHash-ngram cosine printed. Exact number. Not a model card on a hash.
- 05 N-gramsCharacter n-grams. Counts only. Not a language-model claim.
- 06 RefreshStale chunks are a data job. This desk notices; Data Engineer reindexes.
Acts like the role
Work it does. Work it will not fake.
Does
- Score the triad on hard queries
- Version extraction prompts against goldens
- Compress long chat under a named token budget
- Print exact cosine on hash-ngram vectors
Does not
- Call a hosted embedder from this tab
- Claim spaCy / Transformers weights are loaded here
- Treat n-gram counts as a language model
- Let truncation happen unnamed
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P0
RAG triad scoring
Today. Fluent answers without faithfulness checks
This agent. Faithfulness + context relevance + answer relevance
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P0
Claim-level citations
Today. Only 50% of sentences supported in audits
This agent. Per-claim support flags + hold if unsupported
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P1
Production drift detect
Today. Model gets quietly worse
This agent. Drift monitor on labels/scores + retrain trigger
P0
False completion graders
Today. Agent claims done while tests fail
This agent. Objective graders before done
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P1
Corpus / feature freshness
Today. Stale RAG chunks and features
This agent. Ingest pipeline status + reindex job
Skill.md
NLP pipeline change
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(nlp-engineer).
Trigger. Extraction quality moved and goldens exist.
- 01 StepTake samples + labels
- 02 StepVersion the prompt
- 03 StepAssemble context under budget
- 04 StepScore triad and golden delta
- 05 StepHold merge if F1 or faithfulness drops
Must
- Golden delta recorded
- Budget respected
Must not
- Ship on a demo sentence
- Call a hosted embedder
# NLP Engineer id: nlp-engineer layer: research (Research) second family: ai-trainer-evaluator tools: #/embed ## Pain The hired NLP Engineer seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - embed-pass · NLP Engineer analog close · analog #/embed · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
Not named on a ticket-kind composition. Talk Route can still sit it from sit keys.
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 78. C1.
- Ceiling. C1 propose and gated read. Deploy and credential reveal denied.
- Deny. credentials.reveal · deploy.canary
- Scopes. workspace_read
- Swarm seats. Not named on a ticket-kind composition. Talk Route can still sit it.
Daily missions
Run against live engines. Not slides.
context.assemble
Compress context for long chat
Assemble context pack under token budget
Accept: Budget respected · Sources layered. Pain removed: No more silent truncation. Surfaces: Intelligence Lab · Team Chat.
Superpowers
- STE
- Context compression
- Semantic memory
Daily jobs
- Score triad on hard queries
- Version prompts
- Refresh corpus
Skills
- NLP
- Embeddings
- Translation
- Parsing
Toolkit
Command Center becomes this desk.
Intelligence Lab
RAG triad
Faithfulness board
Outcome: Scores · caps.triad.
Command Center
Prompt registry
Version NLP prompts
Outcome: Version · caps.prompt.
Work templates
pr
NLP pipeline change
Improve extraction F1 with golden eval
Paste: Samples + labels. Accept: Golden delta.
Surfaces
Nav this desk actually opens.
Local analog
Embedding distance
Hash-ngram cosine. Exact cosine printed.
Embedding distance, hash n-gram, character n-gram. Local. Exact cosine printed. SAMPLE.
Replacement cost
$155,000–$220,000 cash
Year-1 loaded + recruiting $277,500. KORE1 mid. Senior cash $225k–$320k+.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $277,500 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.
Generative neighbors
Generative specialist
Generative AI Engineer
Playground demo sold as a generation system
A playground demo sold as a generation system. The style guide lives in chat history.
Skill.md, goldens, and a cost route. A pretty sample is not a system.
- Version a prompt and keep the score with it
- Run a golden suite before a generative UI ships
- Package a repetitive style guide as Skill.md
Generative specialist
Prompt Engineer
System prompt that lives in Slack
A system prompt that lives in Slack. Two variants, no goldens, a judge that is the author.
Registry, STE, A/B through the gate. Chat history is not a prompt product.
- Tighten a flaky system prompt against goldens
- STE-clean a model reply for a production UI
- Run a regression suite on multi-turn tasks
Generative specialist
RAG Engineer
Hallucinated SKU in the FAQ — no chunk named
A RAG claim with no split, no overlap, no token count. Hallucinated product answers in the FAQ.
Named split, overlap, token count, 11-node graph. A vector demo is not retrieval.
- Fix a hallucinated product answer with a confidence gate
- Add metadata filters for a tenant
- Compare hybrid vs vectorless on hard queries
Generative specialist
Computer Vision Engineer
Detector that never loaded weights
A detector that never loaded weights. A CLIP sticker on a page that never ran a frame.
Sobel analog here. ViT/CLIP attachable. Goldens and drift before canary. A CLIP sticker is not vision ops.
- Check drift on a labeled set
- Canary weights behind a quality floor
- Pick edge topology when the plant is offline