Foundation specialist · Research

AI Research Scientist

Paper claim that never left the notebook

What this agent actually does

Research → measured prototype.

A paper claim that never leaves the notebook. A diagram with a missing dim treated as compiled.

The pain — unique to this desk

A paper claim that never leaves the notebook. A diagram with a missing dim treated as compiled.

The need this desk closes

A scored experiment card: hypothesis, faithfulness delta, env layer named, lesson written — before anyone swaps the model.

Why a generic chat cannot fake this

Paper claim that never left the notebook

A scored experiment card with env classification. A chatbot summary is not a prototype.

Takes

  • Abstract or lab notes
  • Layer table
  • Baseline description

Returns

  • Hypothesis card
  • Faithfulness delta
  • Lesson in memory

Modalities this desk actually handles

TextLayer tablesPaper figures as topology

How this specialist fuses

Paper claim that never left the notebook

Late fusion on the ticket: topology parse of the figure (not a ViT we host), text ingest of the claim, retrieval bench of the method. Missing dims are flagged — the figure is not compiled.

Vision-language on this desk

Named family. Never assumed.

Attached dual-encoder (CLIP/SigLIP, InfoNCE) ranks figure vs caption. Attached fusion-encoder (BLIP-2 Q-Former, LLaVA MLP, Flamingo gated xattn) captions the figure. Unattached, this desk audits the diagram as a layer table. We do not load those weights here.

Patch / tokenize → project → fuse → ground is the host mechanism. This desk fuses late under its own playbook — not a slogan copied across the roster. Host mechanism.

A note or abstract lands on Talk Route. The desk ingests it into capability proposals, benches hybrid RAG against classic with a faithfulness delta, and treats the 11-method eval gate as the experiment tracker. When a clean env disagrees, harness.classify names the layer instead of swapping the model. Memory consolidates the failure into lessons.md so the next ablation does not repeat it. Deploy and credential reveal stay denied.

Live missions run in Role Studio. Analog kernels stay in this tab until a hosted call is chosen. Authors never grade themselves. A hashed pack is not a letter. Hire bands are market replacement cost, not EngOS payroll. Isolation is probed, not assumed.

Reports up as Research. Second family ai-trainer-evaluator. Authors never grade themselves. Isolation is probed, not assumed.

Five analog ticket kinds

How a hired human spends the week. What the analog kernel closes.

Percent splits stay human-week language. Never GPU-loop quotas. Each kind binds to the analog tool that closes it. Escalate only when the kernel cannot.

  1. AI Research Scientist analog closeHuman bar: Name the ticket. Attach evidence. Second family grades. Analog: #topology. Kernel #/topology on this host. Arithmetic / Canvas 2D. Device none. Escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

A typical ticket

What you say. What sits.

Replicate the retrieval-head claim in this abstract against our FAQ corpus. Record the faithfulness delta and freeze the env.

Sit keys paper / hypothesis / experiment / research / ablation. Ticket kind research. Swarm seats Data Engineer (corpus) and Prompt Engineer (packaging). Eval is a second family — this desk does not grade itself.

Playbook

How this agent works the ticket.

  1. 01 Ingestresearch.ingest parses the note into ranked capability proposals.
  2. 02 Hypothesis cardPaper → experiment template: baseline, claim, acceptance, lesson slot.
  3. 03 Benchmarkrag.grounded vs classic. Both modes scored. Faithfulness delta recorded.
  4. 04 Gateeval.run_gate is merge / hold / reject — 11 methods, not a BLEU spreadsheet.
  5. 05 Isolate envharness.classify when the result differs on a clean env. Do not blame the model first.
  6. 06 RememberEpisodic run consolidates into lessons.md. Same failure next week is a product bug.

Acts like the role

Work it does. Work it will not fake.

Does

  • Reproduce a paper claim with a small prototype and an eval score
  • Compare two retrieval heads on a pasted corpus
  • Classify harness-layer failures without a model swap
  • Log a failed experiment into episodic memory

Does not

  • Ship a canary (deploy.canary denied)
  • Reveal credentials
  • Invent a hosted W&B or JAX runtime on this sales host
  • Treat a paper diagram as compiled dims — topology analog flags missing dims only

Jobs on the board

Scenarios this desk was built to close.

From Role Studio's scenario bank — current pain without Helix, and the job the agent actually runs. Top eight of twenty.

P0

Golden dataset regression

Today. Prompt/model change ships without suite

This agent. Versioned golden set + offline gate before merge

P0

RAG triad scoring

Today. Fluent answers without faithfulness checks

This agent. Faithfulness + context relevance + answer relevance

P1

Episodic → lessons consolidation

Today. Same failures every week

This agent. CoALA memory + procedural skill publish

P0

Prompt git-style versioning

Today. Prompts live in Notion / chat history

This agent. Registry with A/B + golden score

P1

Cost-aware model routing

Today. Single frontier model burns budget

This agent. Token budget + cascade cheap→frontier

P0

Agent loop with durable checkpoints

Today. Agents lose state mid-run; restart from scratch

This agent. Checkpointed harness with resume + budget per cycle

P0

Claim-level citations

Today. Only 50% of sentences supported in audits

This agent. Per-claim support flags + hold if unsupported

P0

False completion graders

Today. Agent claims done while tests fail

This agent. Objective graders before done

Skill.md

Paper → experiment card

Durable analog execution standard. Trigger, five ticket kinds, analog tools, must / must-not, handoff (evidence not authority), spend.halt, second family. Same factory as PR Creator, Code Reviewer, Handoff Verifier — plus fromDesk(ai-research-scientist).

Trigger. A paper, abstract, or lab note needs a scored prototype, not a summary.

  1. 01 StepRead the claim and write the hypothesis in one sentence
  2. 02 StepName the baseline and the metric
  3. 03 StepRun the retrieval or model comparison through the eval gate
  4. 04 StepIf env disagrees, classify the harness layer
  5. 05 StepWrite the lesson even when the claim fails

Must

  • Both modes scored
  • Lesson in memory
  • External evidence only

Must not

  • Grade your own paper
  • Swap the model to hide an env bug
  • Call the pack a letter
# AI Research Scientist

id: ai-research-scientist
layer: research (Research)
second family: ai-trainer-evaluator
tools: #/topology

## Pain
The hired AI Research Scientist seat does not exist yet, or the work has no named ticket.

## Need
A named ticket, an analog close, and a second family. Not a chat that grades itself.

## Isolation
Isolation is probed, not assumed. Unprobed stays unlabeled. Never certified-green from a desk.

## spend.halt
spend.halt on the ticket cap. Only the operator raises the ceiling.

## Ticket kinds
- topology-pass · AI Research Scientist analog close · analog #/topology · escalate when: The analog kernel cannot close, or a hosted connector is required as if it were present.

Who sits with this desk

Swarm compose by ticket kind.

  • research · bar 78

    Data Engineer for AI · Prompt Engineer

    Ingest paper, propose capabilities, benchmark, memory lesson

Sit keys

Talk Route matches these words.

topology

Keyword analog in this tab. Not a hosted model. Ticket id is the idempotency key.

RBAC

Pass bar 78. C2.

  • Ceiling. C2 gated act. Eval bar required. Secrets stay denied unless the profile lists them.
  • Deny. deploy.canary · credentials.reveal
  • Scopes. workspace_read · corpus_read · eval
  • Swarm seats. research

Daily missions

Run against live engines. Not slides.

rag.query_advanced · rag.grounded · eval.run_gate

Benchmark a new retrieval head

Run hybrid RAG vs classic with groundedness claims

Accept: Both modes scored · Faithfulness delta recorded. Pain removed: No more spreadsheet of ad-hoc BLEU scores. Surfaces: Intelligence Lab · Eval Gate.

harness.classify

Freeze model, isolate env failure

Classify harness layer when result differs on clean env

Accept: Layer identified without blaming the model. Pain removed: Stops wasting days on model swaps. Surfaces: Debug Harness.

Superpowers

  • Eval gate as experiment tracker (11 methods)
  • DST single-thread replay for nondeterministic bugs
  • Memory consolidation into lessons.md

Daily jobs

  • Reproduce a paper claim with a small prototype
  • Compare two retrieval heads on your corpus
  • Log a failed experiment into episodic memory

Skills

  • Research
  • Math
  • Deep Learning
  • Experiment design

Toolkit

Command Center becomes this desk.

Command Center

Ingest paper / note

Parse research into capability proposals

Outcome: Ranked features to add to your stack · research.ingest.

Intelligence Lab

Benchmark retrieval

Compare RAG modes with groundedness

Outcome: Faithfulness delta you can cite · rag.grounded.

Eval Gate

Run eval gate

11-method score on candidate output

Outcome: Merge / hold / reject with reasons · eval.run_gate.

Work templates

research

Paper → experiment card

Replicate claim from paper and score against baseline

Paste: Paper abstract or notes + baseline description. Accept: Hypothesis written · Eval scores recorded · Lesson in memory.

Surfaces

Nav this desk actually opens.

Command CenterModels & KeysRole StudioIntelligence LabEval GateAgent MemorySystems LabCollaborative IDECertify & OperateUpdates

Local analog

Paper-to-layer topology

Paste a layer table. Missing dims are flagged. Not a compiler.

Paper-to-layer topology: paste a layer table. Missing dims are flagged. Not a compiler. SAMPLE on this host.

Run Paper-to-layer topology

Replacement cost

$180,000$280,000 cash

Year-1 loaded + recruiting $340,400. KORE1 mid. Senior TC $300k–$489k+ at labs.. Not EngOS payroll. Not ARR.

Pilot $0 / Team $79 sits this analog desk. Year-1 hire is $340,400 loaded. That is not a replacement claim. The hired role remains the real thing. The analog desk reports up so one operator can run the ticket.

Open Talk RouteAll twenty desks

AI EngineerML Engineer

EngOS