The pain — unique to this desk
Foundation specialist · Research
AI Research Scientist
Paper claim that never left the notebook
What this agent actually does
Research → measured prototype.
A paper claim that never leaves the notebook. A diagram with a missing dim treated as compiled.
The need this desk closes
A scored experiment card: hypothesis, faithfulness delta, env layer named, lesson written — before anyone swaps the model.
Why a generic chat cannot fake this
Paper claim that never left the notebook
A scored experiment card with env classification. A chatbot summary is not a prototype.
Takes
- Abstract or lab notes
- Layer table
- Baseline description
Returns
- Hypothesis card
- Faithfulness delta
- Lesson in memory
Modalities this desk actually handles
How this specialist fuses
Paper claim that never left the notebook
Late fusion on the ticket: topology parse of the figure (not a ViT we host), text ingest of the claim, retrieval bench of the method. Missing dims are flagged — the figure is not compiled.
Vision-language on this desk
Named family. Never assumed.
Attached dual-encoder (CLIP/SigLIP, InfoNCE) ranks figure vs caption. Attached fusion-encoder (BLIP-2 Q-Former, LLaVA MLP, Flamingo gated xattn) captions the figure. Unattached, this desk audits the diagram as a layer table. We do not load those weights here.
Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.
A note or abstract lands on Talk Route. The desk ingests it into capability proposals, benches hybrid RAG against classic with a faithfulness delta, and treats the 11-method eval gate as the experiment tracker. When a clean env disagrees, harness.classify names the layer instead of swapping the model. Memory consolidates the failure into lessons.md so the next ablation does not repeat it. Deploy and credential reveal stay denied.
Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.
Reports up as Research. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.
Five analog ticket kinds
How a hired human spends the week. What the analog kernel closes.
Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.
- AI Research Scientist analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #topology. Kernel #/topology on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
A typical ticket
What you say. What sits.
Replicate the retrieval-head claim in this abstract against our FAQ corpus. Record the faithfulness delta and freeze the env.
Sit keys paper / hypothesis / experiment / research / ablation. Ticket kind research. Swarm seats Data Engineer (corpus) and Prompt Engineer (packaging). Eval is a second family — this desk does not grade itself.
Playbook
How this agent works the ticket.
- 01 Ingestresearch.ingest parses the note into ranked capability proposals.
- 02 Hypothesis cardPaper → experiment template: baseline, claim, acceptance, lesson slot.
- 03 Benchmarkrag.grounded vs classic. Both modes scored. Faithfulness delta recorded.
- 04 Gateeval.run_gate is merge / hold / reject — 11 methods, not a BLEU spreadsheet.
- 05 Isolate envharness.classify when the result differs on a clean env. Do not blame the model first.
- 06 RememberEpisodic run consolidates into lessons.md. Same failure next week is a product bug.
Acts like the role
Work it does. Work it will not fake.
Does
- Reproduce a paper claim with a small prototype and an eval score
- Compare two retrieval heads on a pasted corpus
- Classify harness-layer failures without a model swap
- Log a failed experiment into episodic memory
Does not
- Ship a canary (deploy.canary denied)
- Reveal credentials
- Invent a hosted W&B or JAX runtime on this sales host
- Treat a paper diagram as compiled dims — topology analog flags missing dims only
Jobs on the board
Scenarios this desk was built to close.
From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.
P0
Golden dataset regression
Today. Prompt/model change ships without suite
This agent. Versioned golden set + offline gate before merge
P0
RAG triad scoring
Today. Fluent answers without faithfulness checks
This agent. Faithfulness + context relevance + answer relevance
P1
Episodic → lessons consolidation
Today. Same failures every week
This agent. CoALA memory + procedural skill publish
P0
Prompt git-style versioning
Today. Prompts live in Notion / chat history
This agent. Registry with A/B + golden score
P1
Cost-aware model routing
Today. Single frontier model burns budget
This agent. Token budget + cascade cheap→frontier
P0
Agent loop with durable checkpoints
Today. Agents lose state mid-run; restart from scratch
This agent. Checkpointed harness with resume + budget per cycle
P0
Claim-level citations
Today. Only 50% of sentences supported in audits
This agent. Per-claim support flags + hold if unsupported
P0
False completion graders
Today. Agent claims done while tests fail
This agent. Objective graders before done
Skill.md
Paper → experiment card
Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ai-research-scientist).
Trigger. A paper, abstract, or lab note needs a scored prototype, not a summary.
- 01 StepRead the claim and write the hypothesis in one sentence
- 02 StepName the baseline and the metric
- 03 StepRun the retrieval or model comparison through the eval gate
- 04 StepIf env disagrees, classify the harness layer
- 05 StepWrite the lesson even when the claim fails
Must
- Both modes scored
- Lesson in memory
- External evidence only
Must not
- Grade your own paper
- Swap the model to hide an env bug
- Call the pack a letter
# AI Research Scientist id: ai-research-scientist layer: research (Research) second family: ai-trainer-evaluator tools: #/topology ## Pain The hired AI Research Scientist seat does not exist yet, or the work has no named ticket. ## Need A named ticket, an analog close, and a second family. Not a chat that grades itself. ## Isolation Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk. ## spend.halt spend.halt on the ticket cap. Only the operator raises the ceiling. ## Ticket kinds - topology-pass · AI Research Scientist analog close · analog #/topology · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.
Who sits with this desk
Swarm compose by ticket kind.
research · bar 78
Data Engineer for AI · Prompt Engineer
Ingest paper, propose capabilities, benchmark, memory lesson
Sit keys
Talk Route matches these words.
Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.
RBAC
Pass bar 78. C2.
- Ceiling. C2 gated act. Eval bar required. Secrets stay denied unless the profile lists them.
- Deny. deploy.canary · credentials.reveal
- Scopes. workspace_read · corpus_read · eval
- Swarm seats. research
Daily missions
Run against live engines. Not slides.
rag.query_advanced · rag.grounded · eval.run_gate
Benchmark a new retrieval head
Run hybrid RAG vs classic with groundedness claims
Accept: Both modes scored · Faithfulness delta recorded. Pain removed: No more spreadsheet of ad-hoc BLEU scores. Surfaces: Intelligence Lab · Eval Gate.
harness.classify
Freeze model, isolate env failure
Classify harness layer when result differs on clean env
Accept: Layer identified without blaming the model. Pain removed: Stops wasting days on model swaps. Surfaces: Debug Harness.
Superpowers
- Eval gate as experiment tracker (11 methods)
- DST single-thread replay for nondeterministic bugs
- Memory consolidation into lessons.md
Daily jobs
- Reproduce a paper claim with a small prototype
- Compare two retrieval heads on your corpus
- Log a failed experiment into episodic memory
Skills
- Research
- Math
- Deep Learning
- Experiment design
Toolkit
Command Center becomes this desk.
Command Center
Ingest paper / note
Parse research into capability proposals
Outcome: Ranked features to add to your stack · research.ingest.
Intelligence Lab
Benchmark retrieval
Compare RAG modes with groundedness
Outcome: Faithfulness delta you can cite · rag.grounded.
Eval Gate
Run eval gate
11-method score on candidate output
Outcome: Merge / hold / reject with reasons · eval.run_gate.
Work templates
research
Paper → experiment card
Replicate claim from paper and score against baseline
Paste: Paper abstract or notes + baseline description. Accept: Hypothesis written · Eval scores recorded · Lesson in memory.
Surfaces
Nav this desk actually opens.
Local analog
Paper-to-layer topology
Paste a layer table. Missing dims are flagged. Not a compiler.
Paper-to-layer topology: paste a layer table. Missing dims are flagged. Not a compiler. SAMPLE on this host.
Replacement cost
$180,000–$280,000 cash
Year-1 loaded + recruiting $340,400. KORE1 mid. Senior TC $300k–$489k+ at labs.. Not EngOS payroll. Not ARR.
Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $340,400 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.
Foundation neighbors
Foundation specialist
ML Engineer
Notebook that never ships a scorer
A glue notebook that never ships a scorer. A green dashboard on two series nobody compared.
Confusion matrix plus PSI/KS before canary. A notebook pickle is not a model.
- Retrain on last week's labeled tickets and print confusion
- Canary a scorer behind the eval gate
- Profile features and flag drift on pasted series
Foundation specialist
LLM Engineer
Fluent answer with no citation and no fallback
A fluent answer with no citation and no fallback. One frontier call is the whole plan.
Cited claims and a cheap-then-reason chain. Fluency is not a product.
- Ship a grounded FAQ answer a customer can be handed
- Tune the pattern on a failing query
- Add a RAG pattern to the runbook with a scorecard