All in One View

Content from Generative AI as an agentic research tool


Last updated on 2026-09-30 | Edit this page

Overview

Questions

  • How does an agentic assistant differ from a chat model?
  • What are the parts of an LLM “harness”?
  • What loops sit around the agent loop, and what does each add?
  • How do we get trustworthy results from a stochastic model?
  • Why build on an open protocol, not one product?

Objectives

  • Trace one pass of the agent loop for a physics task, naming what enters the context at each step.
  • Draft acceptance criteria precise enough that a grader — human or agent — can apply them.
  • Place a given automation (a format-on-edit hook, a nightly report, a prompt tweak) at the right loop level.
  • Choose between a tool, a skill, and a subagent when adding a new capability.

The bottleneck is rarely the physics


In a typical analysis the physics steps are few: select a final state, build an observable, fit a signal. Most of the time goes into the software around them: finding datasets, learning the data model, getting branch names and units right, and iterating on plotting and analysis code. LLMs can reduce this work, but only if their results can be checked. Later episodes apply this to the decay \(\Lambda^0 \to p\,\pi^-\).

Two modes of use


A chat completion is a single request: a prompt goes in, text comes out. If the text is code, you run it, read the error, and paste it back by hand. The model never observes your data or the result of running anything.

An agentic loop runs the same model inside a control structure that lets it act. The model proposes an action, an external tool carries it out, the result is added to the context, and the model is called again, until a stopping condition is met. The model then works from the actual state of your files and the output of real computations, not only from its training.

Callout

Scope

Here, “generative AI” means an LLM-based coding assistant used in this agentic mode. Machine learning for reconstruction or particle identification is not covered. The subject is writing and running analyses.

This loop is the first of four. Each one wraps the one before it, and this episode goes through them in order.

Level 1 — the agent loop


The basic cycle: the model calls a tool, looks at what came back, and decides whether it is done.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart TD
    accTitle: {Level 1 the agent loop}
    accDescr: {Level 1 the agent loop}
    Q["task"]:::user --> M["model"]:::core
    M -->|"tool call"| T["tools<br/>read files · run code · query "]:::tool
    T -->|"result appended to context"| D{"task<br/>complete?"}:::core
    D -->|"no"| M
    D -->|"yes"| R["result + provenance"]:::out
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
    classDef tool fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
    classDef out fill:#f3e8ff,stroke:#7048e8,stroke-width:1.5px,color:#2e1065;
    classDef user fill:#f1f3f5,stroke:#868e96,stroke-width:1.5px,color:#212529;

The model together with the software that runs this cycle is called a harness. It has four parts:

  • Model: reasoning and code generation. The provider and model can be swapped; they are not part of the method.
  • Context: everything the model sees in a step: instructions, files, earlier turns, tool outputs. Its size is limited (the context window).
  • Tools: operations the model may call. Tools are the only way the model can change anything outside itself.
  • Control loop: repeats propose → execute → observe until the task is done. This is what separates an agent from a chatbot.

Extending the harness

Current assistants add a few standard extension points to this core.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart TB
    accTitle: {AI agent harness}
    accDescr: {AI agent harness}
    MCP["MCP servers<br/>external tools & data"]:::tool --> H
    SUB["subagents<br/>specialized, isolated context"]:::tool --> H
    HOOKS["hooks<br/>lifecycle automation"]:::tool --> H
    MON["monitors<br/>watch & react"]:::tool --> H
    H(["harness core<br/>model + context + tools + loop"]):::core
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
    classDef tool fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
  • MCP servers: tools in a standard form that works with any assistant. This is how the assistant gets physics tools (Episode 3).

  • Subagents: separate assistants started for a sub-task, so a large task can be split up.

  • Skills: versioned procedures (a SKILL.md plus scripts) loaded when a request matches (Episode 4). A tool is a capability; a skill is a recipe.

  • Hooks: scripts the harness runs at fixed moments, e.g. reformat code after every edit or log each tool call.

  • Monitors: watch background work, e.g. a long batch job, and call the assistant again when it finishes.

Level 2 — the verification loop


A level-1 agent stops when it decides the task is complete. An LLM is stochastic: the same prompt can give different outputs, and a confident answer is not evidence of a correct one. A second loop around the first handles this. A grader checks the agent’s result against explicit acceptance criteria, and a failure goes back to the agent as feedback for another attempt.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart TD
    accTitle: {Level 2 the verification loop}
    accDescr: {Level 2 the verification loop}
    A["agent loop  ①"]:::core -->|"draft result"| G["grader<br/>explicit success criteria"]:::tool
    G --> P{"pass?"}:::core
    P -->|"no — feedback into context"| A
    P -->|"yes"| R["accepted result"]:::out
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
    classDef tool fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
    classDef out fill:#f3e8ff,stroke:#7048e8,stroke-width:1.5px,color:#2e1065;

Treat every model output as a hypothesis and accept it only after checking it against something external: the data, a fit statistic, a known physical value, or an independent implementation. For the \(\Lambda^0\) measurement the criteria are physical and checkable: \(|\mu - 1.115683\,\mathrm{GeV}| < 5\,\mathrm{MeV}\), \(\sigma\) consistent with detector resolution, \(\chi^2/\mathrm{ndf}\) of order 1. In Episode 4 you write this grader into the lambda-fit skill’s success criteria. An agent that passes reports the numbers; one that fails must report the failure instead of a result.

Grading costs time and tokens, but in physics analysis correctness matters more than speed. Prefer tools that return small quantities you can inspect (counts, bin edges, fit parameters), and keep a record of what was run.

Level 3 — the operation loop


With verification in place, the agent no longer needs you to start it. In a third loop an event (a schedule, a finished production job, a new pull request) starts the agent, its verified output updates something (a report, a documentation page, an alert), and the system goes back to waiting.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart LR
    accTitle: {Level 3 the operation loop}
    accDescr: {Level 3 the operation loop}
    E["event<br/>schedule · finished job · new PR"]:::pkg --> A["agent + verification<br/>① + ②"]:::core
    A --> U["system update<br/>report · doc page · alert"]:::out
    U -.->|"wait for the next event"| E
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
    classDef out fill:#f3e8ff,stroke:#7048e8,stroke-width:1.5px,color:#2e1065;
    classDef pkg fill:#fff4e0,stroke:#f08c00,stroke-width:1.5px,color:#5c3b00;

Episode 6 shows collaboration examples: documentation regenerated on a schedule, and reviews triggered per pull request.

Level 4 — the improvement loop


The last loop goes around everything. From time to time, look at what the agent actually did (transcripts, failed fits, wrong tool choices) and change the harness accordingly: clarify AGENTS.md, tighten a skill’s success criteria, add a missing tool. You do a small version of this whenever a prompt goes wrong and you fix the instructions instead of retyping the request.

Callout

Where a human belongs in each loop

Each level has a point where a person should decide: approving a sensitive tool call (1), signing off a graded result (2), reviewing what runs unattended (3), and choosing which harness changes to keep (4).

The next episode describes the physics measurement.

Key Points
  • An agentic assistant runs tools and reads their output; a chat model only returns text.
  • Treat every result as a hypothesis and check it against clear criteria before you accept it.

The four-level loop schematics are adapted from The Art of Loop Engineering.

Content from The measurement: Λ⁰ → p π⁻


Last updated on 2026-09-30 | Edit this page

Overview

Questions

  • What is the p π⁻ invariant-mass observable, and its peak and background?
  • Which EDM4eic collections and units does it need?

Objectives

  • Compute m(p, π) by hand from the four momentum branches and the PDG masses.
  • Sketch the expected spectrum (peak position, width scale, background shape) before looking at data.

The Setup page covers the servers and the assistant. This episode describes the physics you reconstruct starting in Episode 3.

The decay


The Λ⁰ is the lightest strange baryon (uds, spin-parity ½⁺). It decays only weakly (a strangeness-changing ΔS = 1 transition), so it is long-lived: cτ ≈ 7.9 cm. Its dominant hadronic mode is

Λ⁰ → p + π⁻      (branching fraction ≈ 63.9%)

The centimeter-scale flight distance makes the decay a V0: two oppositely charged tracks from a vertex displaced from the primary interaction point.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart LR
    accTitle: {Lambda to proton pion V0 decay}
    accDescr: {Lambda to proton pion V0 decay}
    PV["primary vertex<br/>e + A collision"]:::vtx -. "Λ⁰: neutral, cτ ≈ 7.9 cm" .-> DV["displaced<br/>decay vertex"]:::vtx
    DV --> P["proton<br/>PDG 2212"]:::pos
    DV --> PI["pion<br/>PDG -211"]:::neg
    classDef vtx fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
    classDef pos fill:#ffe3e3,stroke:#e03131,stroke-width:1.5px,color:#5c0a0a;
    classDef neg fill:#e7f5ff,stroke:#1971c2,stroke-width:1.5px,color:#0a3d62;

The observable

The Λ⁰ is neutral and not detected directly; we reconstruct it from its charged daughters. For a candidate proton \(p_1 = (E_1, \vec{p}_1)\) and candidate pion \(p_2 = (E_2, \vec{p}_2)\), the pair’s invariant mass is Lorentz invariant:

\[E_i = \sqrt{|\vec{p}_i|^2 + m_i^2}\]

with \(m_i\) the assigned proton or pion mass, and

\[m(p, \pi) = \sqrt{(E_1 + E_2)^2 - |\vec{p}_1 + \vec{p}_2|^2}\]

Assign the proton mass to one track and the pion mass to the other (using reconstructed particle ID). For true Λ⁰ decays this equals the parent mass; candidates accumulate in a peak at 1.115683 GeV.

Callout

Width: resolution, not lifetime

The Λ⁰ natural width (\(\Gamma = \hbar/\tau \approx 2.5 \times 10^{-6}\) eV) is far below any detector effect. The observed peak width, a few MeV, measures the detector momentum and angular resolution, not the particle.

Background

Most proton–pion pairs do not come from a Λ⁰ at all. These random (“combinatorial”) pairs do not peak; they form a smooth distribution under the signal. The analysis extracts a yield by fitting a Gaussian peak on top of a low-order polynomial background (Episode 5). The charge-conjugate mode Λ̄ → p̄ π⁺ is reconstructed identically with the antiparticles.

Callout

Reference values (PDG)

Quantity Value
m(Λ⁰) 1.115683 GeV
m(p) 0.9382720813 GeV
m(π±) 0.13957061 GeV
cτ(Λ⁰) 7.89 cm
BR(Λ⁰ → p π⁻) 63.9 %

Energies and momenta are in GeV (natural units, c = 1).

The data model


ePIC reconstruction output uses EDM4eic, an EIC extension of EDM4hep generated with PODIO. A file contains an events tree; each entry is one event, each branch a collection. We need one collection, the reconstructed charged tracks, and four members:

events  (tree; one entry per event)
    ReconstructedChargedParticles.PDG          reconstructed particle-ID hypothesis
    ReconstructedChargedParticles.momentum.x   p_x  [GeV]
    ReconstructedChargedParticles.momentum.y   p_y  [GeV]
    ReconstructedChargedParticles.momentum.z   p_z  [GeV]

PDG is the Particle Data Group code the reconstruction assigns each track. Select protons (2212) and π⁻ (-211) for Λ⁰, antiprotons (-2212) and π⁺ (211) for Λ̄.

You do not download a file. In Episode 3 the assistant uses the rucio tools to find a DIS dataset and xrootd to verify its files, then reads one of the dataset’s root:// URLs (e.g. root://epicxrd1.sdcc.bnl.gov:1095//...) in place with the uproot tools. It reads these branches without you writing any I/O code.

Key Points
  • The Λ⁰ shows up as a peak at 1.1157 GeV in the proton–pion invariant mass.
  • The peak width comes from detector resolution.

Content from Tool servers and the Model Context Protocol (MCP)


Last updated on 2026-09-30 | Edit this page

Overview

Questions

  • What is MCP, and what problem does it solve?
  • What can the EIC tool servers do?
  • How do you connect an assistant to them?

Objectives

  • Start the servers, check them, and read a server log when something fails (eic-mcp up/status/logs).
  • Generate the connection file for your own client with eic-mcp config.
  • Discover a real DIS dataset by prompting, without hard-coding names or paths.
  • Judge which returned quantities are worth verifying, and against what.

One interface for tools


Tools are the only way an assistant can act (Episode 1). The Model Context Protocol (MCP) defines a standard interface for them: write a tool once as a server, and any client (assistant) that supports MCP can use it.

The lesson’s servers run inside eic-shell, and the assistant talks to them over a local web address.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart LR
    accTitle: {EIC MCP data tools}
    accDescr: {EIC MCP data tools}
    A["AI assistant<br/>opencode · Copilot · Cursor"]:::core <-->|"MCP"| S["uproot tool server<br/>(MCP, in eic-shell)"]:::tool
    S <-->|"uproot"| F["EDM4eic ROOT file"]:::data
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
    classDef tool fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
    classDef data fill:#fff4e0,stroke:#f08c00,stroke-width:1.5px,color:#5c3b00;

The uproot tool server


The ePIC uproot tool server reads ROOT/EDM4eic files with uproot. It can list a file’s contents, compute statistics and histograms, and run short NumPy calculations over one file or a whole dataset. It returns small summaries (counts, bin edges, statistics) that you can check, not raw data. The calculations run in a sandbox that cannot install software or write files.

Start the servers


Start the servers inside eic-shell (see Setup):

BASH

$ eic-mcp up

This starts the uproot, xrootd, and rucio servers. eic-mcp status shows which are running, and eic-mcp logs uproot shows a server’s log.

Callout

If rucio answers but xrootd/uproot time out

If dataset queries work but file access hangs, the XRootD store may be down; check with xrdfs root://epicxrd1.sdcc.bnl.gov:1095 ls /eic/EPIC/RECO.

The default store (BNL) has campaigns from 25.12.0. For older ones (up to 25.10.x) use JLab: XROOTD_SERVER=root://dtn2304.jlab.org:8443 XROOTD_BASE_DIR=/jlab-osdf-ro/eic/EPIC/volatile eic-mcp restart.

If every uproot call times out after one large call, the server is busy, not broken: it handles one request at a time. Wait, or run EIC_MCP_SERVERS=uproot eic-mcp restart.

Connect the assistant


Write opencode’s config file in the directory where you start opencode, then start it:

BASH

$ eic-mcp config opencode
$ opencode

In the session, /mcp lists the connected servers and their tools.

Other clients use the same URLs: eic-mcp config claude (or copilot, vscode, cursor, gemini, codex) writes the file where that client reads it.

Finding the data with MCP


You do not download a dataset. The other two MCP servers let the assistant find and verify the files, in place of running rucio and xrdfs by hand, and uproot-mcp then reads them directly from the store.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart LR
    accTitle: {EIC MCP data tools}
    accDescr: {EIC MCP data tools}
    R["rucio-mcp<br/>find the dataset"]:::tool -->|"file locations"| X["xrootd-mcp<br/>check the files"]:::tool
    X -->|"checked files"| U["uproot-mcp<br/>analyze in place"]:::core
    classDef tool fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
  • rucio-mcp searches the data catalog: it finds a dataset by name, lists its files, and gives their root:// locations.
  • xrootd-mcp browses the data store and checks that the files exist.

uproot-mcp then reads a root:// file in place.

List the available campaigns


ePIC data is organized by production campaign, a version such as 26.06.0. The campaign, the beam/target, and the physics process are all part of the rucio DID (e.g. epic:/RECO/26.06.0/epic_craterlake/DIS/pythia8.316-1.0/NC/noRad/ep/18x275/...). Before locating a specific dataset, check which campaigns exist so you use a current one:

Using the rucio tools, find which production campaigns are available (the version field in the DIDs, e.g. 26.06.0) and show the most recent few.

Watch how the assistant does this: the catalog holds thousands of datasets in no particular order, so looking at only the first page of results can miss the newest campaigns.

Challenge

Exercise: locate a dataset (≈ 10 min)

Ask your assistant:

Use the rucio tools to find the ePIC reconstructed-DIS dataset for the BeAGLE eCu ep 10x115 GeV sample in campaign 26.04.1, list its files, then use the xrootd tools to confirm those files exist on the store and report the total number of events.

The assistant finds the dataset with rucio (374 files), gets their root:// locations, and checks them with xrootd. rucio does not store event counts, and reading all 374 files would take an hour, so a good answer checks a few files (≈ 1,220 events each) and extrapolates.

Inspect the dataset


You describe what you want in plain language and the assistant makes the tool calls. Take one of the root:// URLs from the previous exercise (written below as root://epicxrd1.sdcc.bnl.gov:1095//…) and analyze it in place.

Challenge

Exercise: enumerate the schema (≈ 10 min)

Issue the request:

Using the uproot tools, report the structure of the events tree in root://epicxrd1.sdcc.bnl.gov:1095//<your-discovered-file>.root and list the members of the ReconstructedChargedParticles collection.

The assistant reads the structure of the events tree and reports something like:

OUTPUT

File:  root://epicxrd1.sdcc.bnl.gov:1095//…/<dataset-file>.root
Tree:  events   — branches grouped by collection

ReconstructedChargedParticles collection:
  ReconstructedChargedParticles.PDG          int32[]   PDG particle-ID code
  ReconstructedChargedParticles.momentum.x   float[]   p_x [GeV]
  ReconstructedChargedParticles.momentum.y   float[]   p_y [GeV]
  ReconstructedChargedParticles.momentum.z   float[]   p_z [GeV]
  … energy, charge, mass, type, referencePoint.*, covMatrix.*

The names are read from the file, not guessed, so the assistant cannot invent branch names.

Challenge

Exercise: identify the species present (≈ 10 min)

Issue the request:

Histogram ReconstructedChargedParticles.PDG with one bin per integer code, so I can see the reconstructed particle species in the file.

The assistant makes a histogram with one bin per PDG code. A reconstructed-DIS file gives, for example:

OUTPUT

   PDG  species   count
  -211   pi-      11447
   211   pi+       9885
    11   e-        4489
     0   unID      2971      <- tracks with no PID hypothesis
  -321   K-        1662
   321   K+        1588
   -11   e+         967
 -2212   pbar       693
  2212   p          684      <- protons are rare
Bar histogram of reconstructed charged-particle PDG codes in the file, with pions dominating and protons rare
Reconstructed charged-particle species in the file

Pions dominate; protons are rare (≈ 2%), so the Λ⁰ signal will be small. A sizeable fraction of tracks have no PID (code 0) or a wrong one. This misidentification adds to the combinatorial background, which is why we fit the peak instead of counting it.

Callout

Verify the returned quantities

Look at the returned numbers (bin edges, counts, statistics). Do the PDG peaks fall at physical codes, and are the proton and pion yields plausible? Episode 4 turns this into explicit success criteria.

The same Λ⁰ peak can be obtained without MCP, with ROOT RDataFrame, TTreeReader, plain uproot, or the PODIO Frame API; scripts are in extras/.

The assistant can now query the data through tools whose output you can check. The next episode writes this procedure down as a reusable, versioned skill.

Key Points
  • MCP servers give an assistant tools; any MCP assistant can use them.
  • eic-mcp up starts the servers in eic-shell and eic-mcp config opencode connects opencode.

Content from Persisting instructions: AGENTS.md and SKILL.md


Last updated on 2026-09-30 | Edit this page

Overview

Questions

  • How do you give an assistant durable project context?
  • AGENTS.md or SKILL.md: when to use each?
  • How do you make every tool read the same rules?
  • What’s in a usable SKILL.md for the Λ⁰ fit?

Objectives

  • Write an AGENTS.md for always-on project context.
  • Bridge it so one AGENTS.md drives any tool.
  • Write a SKILL.md that runs the Λ⁰ fit on demand.
  • Encode success criteria + provenance for auditable output.

Two ways to make instructions persistent


Typed requests (Episode 3) have to repeat the data model, conventions, and procedure every session, and two runs can differ. Two kinds of file solve this.

  • AGENTS.md: context read at the start of every session: environment, data model, conventions, and what “done” means.
  • SKILL.md: a procedure in a named skill directory, loaded only when a request matches its description. It describes one repeatable workflow.
%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart TD
    accTitle: {Skills and AGENTS.md}
    accDescr: {Skills and AGENTS.md}
    R["your project"] --> AG["AGENTS.md<br/>whole file always in context"]:::always
    R --> SK[".opencode/skills/lambda-fit/SKILL.md<br/>only its description is indexed"]:::ondemand
    AG --> M(["model context"]):::core
    SK -. "body loaded only when a<br/>request matches its description" .-> M
    classDef always fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
    classDef ondemand fill:#fff4e0,stroke:#f08c00,stroke-width:1.5px,color:#5c3b00;
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;

AGENTS.md answers “what is this project and how do we work here?”; a SKILL.md answers “how do I carry out this task?”.

AGENTS.md: project context


AGENTS.md is plain Markdown at your project root (subdirectories may override it for files beneath them). opencode, Codex, Gemini CLI, Zed, and others read it automatically; a tool that looks for a different filename reads the same content through a one-line bridge (next section).

It is loaded on every turn, so keep it short and factual:

MARKDOWN

# AGENTS.md — Lambda analysis project

## What this project does
Reconstruct Lambda0 -> p pi- in ePIC EDM4eic data and fit the invariant-mass
peak near 1.115683 GeV.

## Environment
- Everything runs inside eic-shell; the MCP servers are started with `eic-mcp up`.
- Data lives on the grid: find a DIS dataset with the `rucio` tools and read its
  root:// files in place with `uproot`. Do not download.

## Tools
- Use the `rucio` MCP server (list_dids, list_files, list_file_replicas) to locate
  a dataset and resolve its root:// URLs.
- Use the `xrootd` MCP server (check_file_exists, get_file_info) to verify a file.
- Use the `uproot` MCP server (get_tree_info, histogram_branch, execute_kernel,
  execute_kernel_dataset) for all ROOT file access. Prefer get_tree_info over
  get_file_structure: on an EDM4eic file the latter returns megabytes.
- Do NOT write bespoke file I/O; the servers already handle it.

## When a tool fails

- Never install software (no `pip install`, above all not `--break-system-packages`)
  and never re-implement the analysis with local uproot/ROOT.
- A timed-out call means the server is BUSY, not broken: it is single-threaded and
  still working on the previous request. Wait, retry once, and if it still fails,
  stop and report which tool failed with which arguments.
- Never reuse a cached earlier tool result as if it were fresh. A number that did
  not come from the MCP servers is not reproducible, so it is not an answer.

## Data model
- Tree: events.  Collection: ReconstructedChargedParticles.
- Members: .PDG, .momentum.x, .momentum.y, .momentum.z   (momenta in GeV).
- PDG codes: proton 2212, pi- -211, antiproton -2212, pi+ 211.

## Physics constants (PDG)
- m(proton) = 0.9382720813 GeV, m(pi) = 0.13957061 GeV, m(Lambda) = 1.115683 GeV.

## Conventions
- Invariant mass over [1.05, 1.25] GeV, 200 bins.
- Fit a Gaussian + 2nd-order polynomial over [1.08, 1.16] GeV.
- Write results as JSON; save plots under output/.

## Definition of done
- Fitted peak within a few MeV of 1.115683 GeV and chi2/ndf of order 1.
- Always run the fit and check these before reporting a result.

It records the schema, the tool policy (use the server, not hand-written I/O), the conventions, and a definition of done.

Callout

Two rules for a useful AGENTS.md

  • Keep it short. It is loaded on every turn, so every line costs tokens. Write only what the model cannot infer from the code.
  • Describe concepts, not file paths. “The reconstructed tracks are in the ReconstructedChargedParticles collection” stays true; a path like src/old/lambda_v2.py goes stale, and the model then searches in the wrong place.

One source of truth: bridge files


opencode, Codex, Gemini CLI, and Zed read AGENTS.md. Some tools read only their own file: GitHub Copilot reads copilot-instructions.md, Cursor reads .cursorrules. Without it they run with no context and no warning.

Do not copy your rules into a second file; the copies will diverge. Make the tool-specific file a one-line bridge to AGENTS.md. .github/copilot-instructions.md (Copilot) and .cursorrules (Cursor) contain:

MARKDOWN

Follow the project rules in AGENTS.md.

The bridge file is in the repository: copilot-instructions.md.

SKILL.md — a named procedure


A skill is a directory containing a SKILL.md that describes a repeatable procedure:

BASH

skills/
  lambda-fit/
    SKILL.md              specification: applicability, inputs, steps, success criteria

The procedure only uses the MCP tools, so it needs no scripts of its own.

The YAML frontmatter has a name and a description. The client matches your request against the description to decide whether to load the skill. Only the name and description stay in context; the body is read only when the description matches.

MARKDOWN

---
name: lambda-fit
description: >
  Reconstruct and fit the Lambda0 -> p pi- invariant-mass peak in ePIC EDM4eic
  data. Use when asked to measure the Lambda yield, mass, or width, or to
  reproduce the Lambda peak from a .root file or a file list.
---

# Lambda invariant-mass fit

## When to use
Any request to find, fit, or quantify the Lambda0 (or its antiparticle) in ePIC
reconstructed data via the proton-pion invariant mass.

## Inputs
- file: one EDM4eic .root URL (a root:// file from a DIS dataset), or
- file_list: the dataset's root:// files for the full sample
  (resolve both with the rucio tools: list_dids, list_files, list_file_replicas).

## Steps
1. Confirm the uproot MCP server is connected: get_tree_info on the input.
2. Build the proton-pion invariant-mass histogram with execute_kernel (one file)
   or execute_kernel_dataset (many files), tree_name 'events' and the
   ReconstructedChargedParticles momentum/PDG branches. For a large sample, cap
   the file count first; for more than ~10 files use submit_kernel_dataset and
   poll, in batches of ~20 files per job (one big job can hit an upstream idle
   timeout), so no single tool call outlives the client's timeout. Write any
   reduce/merge code as plain NumPy array operations (the sandbox rejects tuple
   unpacking in loops).
3. Fit the histogram with a second execute_kernel call (Gaussian + 2nd-order
   polynomial over [1.08, 1.16] GeV; NumPy/awkward only, no imports).
4. Report mu, sigma, signal yield S, and chi2/ndf.

## Success criteria (check before reporting success)
- At least ~50 entries in the fit window. With fewer, report insufficient
  statistics and stop — a low-stats fit lets the polynomial absorb the peak and
  can pass the checks below by accident.
- |mu - 1.115683 GeV| < 0.005 GeV.
- sigma in ~[0.001, 0.005] GeV (this is detector resolution, not natural width).
- chi2/ndf of order 1.
If any check fails, report the failure and the fit diagnostics, not a result.

## Provenance
List the tool calls and their parameters, and the dataset used (campaign and
file list), so the run can be reproduced.
Callout

How clients load a skill

opencode reads skills from .opencode/skills/<name>/SKILL.md in the project directory (or ~/.config/opencode/skills/ for all projects); Claude Code uses .claude/skills/. Download the lesson’s copy:

BASH

curl -fsSL --create-dirs -o .opencode/skills/lambda-fit/SKILL.md \
  https://raw.githubusercontent.com/eic/tutorial-mcp/main/files/skills/lambda-fit/SKILL.md

Loading is the model’s decision, based on the skill’s description. A small model may answer without loading it, so every prompt in this lesson names the skill: “Using the lambda-fit skill”. Clients without a skill mechanism can reference the procedure from AGENTS.md instead.

Get both example files in place: files/skills/AGENTS.md (download it to your analysis directory) and files/skills/lambda-fit/SKILL.md (download it as above).

Your project layout


A project that behaves the same under any assistant:

BASH

lambda-analysis/
├── AGENTS.md                        # source of truth: context + conventions (write this)
├── .github/
│   └── copilot-instructions.md      # points to AGENTS.md   (bridge for Copilot)
├── .cursorrules                     # points to AGENTS.md   (bridge for Cursor)
├── opencode.jsonc                   # MCP server connections: `eic-mcp config opencode` (Episode 3)
└── .opencode/
    └── skills/
        └── lambda-fit/              # downloaded from the lesson (see callout above;
            └── SKILL.md             #  `.claude/skills/` for Claude Code)
Challenge

Exercise: a summary skill (≈ 10 min)

Write a minimal SKILL.md for “summarize the contents of any EDM4eic file”, place it where your client loads skills, and try it on a file from Episode 3.

MARKDOWN

---
name: edm4eic-summary
description: >
  Summarize the contents of an EDM4eic .root file. Use when asked what a
  reconstruction file contains, which trees or collections it holds, or
  how many events it has.
---

# EDM4eic file summary

## Steps
1. get_tree_info on the `events` tree: entry count and collection names.
   (Skip get_file_structure — on EDM4eic files it returns megabytes.)
2. get_tree_info on `runs` and `podio_metadata` for provenance.
3. Return a compact summary: each tree with its entry count, and the
   top-level collections grouped by kind (truth, tracking, calorimetry,
   PID, reconstructed).

## Success criteria
Every tree named with its entry count; if a tree is missing, say so
rather than guessing.

Save it as .opencode/skills/edm4eic-summary/SKILL.md and name it in the prompt (“Using the edm4eic-summary skill, …”); a small model may not load it from the description alone.

Challenge

Exercise: richer provenance (≈ 5 min)

Extend the provenance section of lambda-fit so a run also records the number of input files and the total number of candidate pairs.

Replace the skill’s Provenance section with:

MARKDOWN

## Provenance
List the tool calls and their parameters, the dataset used (campaign and
file list), the number of input files processed, and the total number of
proton-pion candidate pairs entering the histogram, so the run can be
reproduced.

The kernel can return the pair count next to the histogram (e.g. {"counts": ..., "n_pairs": int(len(m))}) and sum it over files.

The next episode runs this skill end to end and scales it from one file to the full sample.

Key Points
  • AGENTS.md holds project context; a SKILL.md holds a procedure the assistant loads when needed.
  • Put success criteria in the skill so you can check the result.

Content from An end-to-end, reproducible Λ⁰ analysis


Last updated on 2026-09-30 | Edit this page

Overview

Questions

  • How do assistant + server + skill compose into one analysis?
  • How does the kernel scale from one file to the full sample?
  • How is the yield extracted and made reproducible?

Objectives

  • Set up any MCP client for the end-to-end run in three steps.
  • Scale the analysis from one file to a full dataset.
  • Accept or reject an agent-produced yield using the audit checklist.

Run it in your client


In eic-shell, in your analysis directory:

BASH

eic-mcp up
eic-mcp config opencode
R=https://raw.githubusercontent.com/eic/tutorial-mcp/main/files/skills
curl -fsSLO $R/AGENTS.md
curl -fsSL --create-dirs -o .opencode/skills/lambda-fit/SKILL.md $R/lambda-fit/SKILL.md
opencode

Check that /mcp lists uproot, xrootd, and rucio, then paste the prompts below. For Claude Code, use eic-mcp config claude and .claude/skills/.

The composed pipeline


The skill (Episode 4) gives the steps, the uproot server (Episode 3) reads the data, and the agent loop (Episode 1) runs and checks it.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart LR
    accTitle: {End-to-end agent run}
    accDescr: {End-to-end agent run}
    A["resolve input<br/>root:// file or file list"]:::data --> B["build m(p,π) histogram<br/>uproot MCP"]:::tool
    B --> C["fit Gaussian + poly-2<br/>opencode prompt"]:::tool
    C --> D["report μ, σ, S, χ²/ndf<br/>+ plot + provenance"]:::out
    classDef data fill:#fff4e0,stroke:#f08c00,stroke-width:1.5px,color:#5c3b00;
    classDef tool fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
    classDef out fill:#f3e8ff,stroke:#7048e8,stroke-width:1.5px,color:#2e1065;

One file, end to end


With the three servers running and the lambda-fit skill available, one request runs the whole chain. Point it at one of the dataset’s root:// files. The assistant uses rucio tools to find a DIS dataset and list_file_replicas for the URLs, xrootd to confirm the file is there, then reads it in place:

Using the lambda-fit skill, measure the Lambda0 peak in this file:
root://epicxrd1.sdcc.bnl.gov:1095//... (one root:// file of the campaign 26.04.1
BeAGLE eCu ep 10x115 dataset from Episode 3's exercise).
Build the proton-pion invariant-mass histogram with the uproot MCP server (tree 'events'),
fit it, and report mu, sigma, the yield, and chi2/ndf, with the plot, or report insufficient
statistics if the fit window is too sparse.

The assistant builds the histogram with the uproot server, then fits it. If your model stops after the histogram (small models do), paste the fit as its own request:

Fit the histogram you just built with execute_kernel: a Gaussian plus a 2nd-order polynomial
over [1.08, 1.16] GeV. Report mu, sigma, the signal yield S, and chi2/ndf, checked against
the lambda-fit skill's success criteria.

One file gives only ~8 candidates in the fit window, so the correct answer is insufficient statistics: the skill says to stop instead of fitting, and a model that quotes \(\mu\) anyway has ignored it. The peak at \(\mu \approx 1.1157\) GeV comes from the larger sample below.

Callout

Smaller models take shortcuts: check the result, not the route

A smaller model may call different tools than the skill says and still get the same histogram. The audit checklist below judges the result, not the route.

Scaling to the full sample


The same calculation can run over many files and return one merged histogram:

Using the lambda-fit skill, run the same proton-pion mass kernel across the first 8 files of the
campaign 26.04.1 BeAGLE eCu ep 10x115 dataset with execute_kernel_dataset (tree 'events'),
merge the histograms, then fit the result and report mu, sigma, the yield, and chi2/ndf
for both Lambda and anti-Lambda, with the plot.

Eight files take about a minute (paste the fit prompt above if it stops). For more files, run background jobs in batches of ~20 files and let the assistant wait for them:

Using the lambda-fit skill, run the proton-pion mass kernel over the first 100 files of the
campaign 26.04.1 BeAGLE eCu ep 10x115 dataset. Submit it with submit_kernel_dataset
(tree 'events') in batches of about 20 files, poll get_job_status until every job finishes,
fetch the histograms with get_job_result and merge them, then fit and report mu, sigma,
the yield, and chi2/ndf for both Lambda and anti-Lambda, with the plot.

Over ~100 files this gives the reference spectrum below: a clear Λ⁰ (and Λ̄) peak over the combinatorial background.

Proton–pion invariant-mass spectrum with Gaussian-plus-polynomial fits showing clear Lambda and anti-Lambda peaks near 1.1157 GeV
Fitted Λ⁰ and Λ̄ invariant-mass spectra (reference fit)

OUTPUT

Lambda      -> p pi-:   mu = 1116.30 +/- 0.32 MeV   sigma = 2.72 +/- 0.33 MeV   S = 123   chi2/ndf = 1.16
anti-Lambda -> pbar pi+: mu = 1116.06 +/- 0.33 MeV   sigma = 3.35 +/- 0.34 MeV   S = 160   chi2/ndf = 1.05

The fitted \(\mu\) sits ~0.6 MeV above the PDG value (1.115683 GeV), a calibration-level offset typical of reconstructed momenta; \(\sigma\) is the detector mass resolution, not the (negligible) Λ⁰ natural width.

Extracting the yield


The fit model is a Gaussian signal on a second-order polynomial background over \([1.08, 1.16]\) GeV:

\[ f(m) = A \exp\!\left[ -\tfrac{1}{2} (m - \mu)^2 / \sigma^2 \right] + \left( c_0 + c_1 (m - m_\Lambda) + c_2 (m - m_\Lambda)^2 \right) \]

The polynomial absorbs the combinatorial background (Episode 2); the integrated signal is \(S = A\sqrt{2\pi}\,\sigma / (\text{bin width})\). Report \(S\) with its uncertainty alongside \(\mu\), \(\sigma\), and \(\chi^2/\text{ndf}\); a bare bin count mixes signal with background.

Audit


Callout

Audit checklist

  • Signal. \(\mu\) within a few MeV of 1.115683 GeV; \(\sigma\) consistent with detector resolution; \(\chi^2/\text{ndf}\) of order unity; \(S\) reported with an uncertainty.
  • Inputs pinned. Dataset (campaign and file list), particle masses, mass window, binning, and fit range all fixed and recorded.
  • Provenance. Tool calls and their arguments logged, so the run can be reconstructed.
  • Cost bounded. A few files during development before scaling up.
  • Oversight. A human inspected the fit before the result was reported.

Exercises (specification)


  • Run the single-file chain through your assistant and report \(\mu\), \(\sigma\), \(S\), and \(\chi^2/\text{ndf}\).
  • Process 8 files and compare the fit to the ~100-file result; comment on the statistical uncertainty.
  • Complete the audit checklist for your run, attaching the recorded tool calls as provenance.
Callout

Try a different beam

The same prompt works with any reconstructed-DIS DID. Nuclear beams (eCu, eAu) give the most Λ per event; ep samples need a few times more files.

The final episode lists the other MCP servers the EIC provides.

Key Points
  • One prompt runs the whole Λ⁰ analysis, from one file up to the full dataset.
  • Check the fit against the audit checklist before you trust it.

Content from Catalog: MCP servers and AI infrastructure in the EIC ecosystem


Last updated on 2026-09-30 | Edit this page

Overview

Questions

  • Which other MCP servers does EIC provide?
  • How can you use them without any setup?

Objectives

  • Ask the DISpatcher bot a question about data, software, or production.

More EIC servers


The three servers you used are part of a larger set, built mostly in BNL’s NPPS group (this episode is based on Torre Wenaus’s June 2026 talk to the ePIC user-learning WG). The eic GitHub organization has the current list.

%%{init: {'theme':'base', 'themeVariables': {'fontSize':'15px','lineColor':'#94a3b8','edgeLabelBackground':'#e2e8f0','clusterBkg':'#1f293720','clusterBorder':'#94a3b8','titleColor':'#94a3b8'}}}%%
flowchart TB
    accTitle: {EIC MCP server catalog}
    accDescr: {EIC MCP server catalog}
    A(["your AI assistant"]):::core
    A --> DATA
    A --> REC
    A --> CODE
    A --> PROD
    subgraph DATA["analysis & data"]
        direction LR
        UP["uproot-mcp"]:::tool
        XR["xrootd-mcp"]:::tool
        RU["rucio-mcp"]:::tool
    end
    subgraph REC["records"]
        direction LR
        ZE["zenodo-mcp"]:::rec
    end
    subgraph CODE["code knowledge"]
        direction LR
        LX["LXR-mcp · BNL-hosted"]:::code
        GH["GitHub-mcp · standard"]:::code
    end
    subgraph PROD["production · via the bot"]
        direction LR
        PB["PanDA · PCS · streaming"]:::pkg
    end
    classDef core fill:#e7efff,stroke:#4c6ef5,stroke-width:1.5px,color:#10204a;
    classDef tool fill:#e6f7ed,stroke:#2f9e44,stroke-width:1.5px,color:#0b3d1f;
    classDef rec fill:#f3e8ff,stroke:#7048e8,stroke-width:1.5px,color:#2e1065;
    classDef code fill:#fff4e0,stroke:#f08c00,stroke-width:1.5px,color:#5c3b00;
    classDef pkg fill:#ffe3e3,stroke:#e03131,stroke-width:1.5px,color:#5c0a0a;
    click UP "https://github.com/eic/uproot-mcp-server" _blank
    click XR "https://github.com/eic/xrootd-mcp-server" _blank
    click RU "https://github.com/eic/rucio-eic-mcp-server" _blank
    click ZE "https://github.com/eic/zenodo-mcp-server" _blank
    click LX "https://eic-code-browser.sdcc.bnl.gov/lxr/source" _blank
    click GH "https://github.com/github/github-mcp-server" _blank
    click PB "https://chat.epic-eic.org/main/channels/dispatcher" _blank

Click a box to open the server’s repository or page.

Analysis and data


Callout

uproot-mcp: read ROOT/EDM4eic files · available · used in this lesson

uproot logo

eic/uproot-mcp-server reads ROOT files. Used in Episodes 3 and 5.

Callout

xrootd-mcp: find files on the data store · available · used in this lesson

XRootD logo

eic/xrootd-mcp-server browses the ePIC data stores.

Callout

rucio-mcp: query the data-management system · available · used in this lesson

Rucio logo

eic/rucio-eic-mcp-server searches the Rucio data catalog (read-only).

Records


Callout

zenodo-mcp: search the open-data repository · available

Zenodo logo

eic/zenodo-mcp-server searches Zenodo, including ePIC documents.

Code knowledge


Callout

LXR-mcp: source cross-reference · available (BNL-hosted)

Searches the ePIC source code through the LXR browser, which is updated nightly. It runs only on the BNL-hosted services.

No setup: the DISpatcher bot


DISpatcher is a Mattermost bot (chat.epic-eic.org → dispatcher) that anyone in ePIC can use, in the channel or by DM. It has the data tools from this lesson plus tools for production jobs, physics samples, software, and documents.

Post this in the dispatcher channel or DM the bot. Your own assistant has no PCS tool and would have to invent the answer:

Summarize the physics tags in the PCS: which processes are covered, and which tags are still draft?

The EIC software portal also has an AI search box (“Ask anything about EIC…”).

corun-ai


BNLNPPS/corun-ai runs longer jobs with larger models and keeps the results. Its first use, codoc-ai, writes software documentation and reviews ePIC pull requests. Ask Torre for an account.

Key Points
  • EIC provides more MCP servers than the three used here.
  • The DISpatcher bot in Mattermost gives you these tools without any setup.