← SELECTED WORK

Live nightly index · human-labelled accuracy pending

MCP Observatory

A source-only MCP scanner with disclosure enforced in code.

Problem
MCP plugins can access files and shell. Review should not require running them.
Decision
Read source without executing servers, and enforce disclosure at the output boundary.
Result
1,673 scanned · 171 findings published

2026-10-04 snapshot. Accuracy is unmeasured; findings are not confirmed vulnerabilities.

Problem

MCP servers connect AI assistants to tools such as shell commands and filesystem access. I wanted to inspect those plugins without running untrusted software, while keeping suspicious source patterns separate from verified vulnerabilities. That distinction is why the scanner’s eventual accuracy measurement matters as much as its finding count.

My contribution

I authored the corpus construction, static rules, sampling design, disclosure boundary, nightly pipeline and public dashboard. I also built the triage and evaluation machinery, while keeping its unwired state explicit.

Constraints

  • Read source only: no dependency installation, build step or execution of a scanned server.
  • Count repositories rather than registry names; a reproducible sample must identify its corpus, seed and date.
  • Withhold high and critical findings until their disclosure window has closed; missing notification records fail closed.
  • Keep language coverage and unmeasured accuracy visible beside the findings.

Architecture & consequential decisions

  1. Sample2,000 repositoriesSeeded, reproducible
  2. Inspect source1,673 servers scannedNo server execution
  3. Gate publication171 public · 745 withheld916 candidate findings
2026-10-04 snapshot. Findings are candidates, not confirmed vulnerabilities. Withholding is a disclosure boundary; accuracy remains unmeasured.

CORPUS → SAMPLE → READ-ONLY FETCH

The crawler removes hosted entries with no source and collapses names pointing to the same repository. The 2026-10-04 aggregate records a 23,157-repository corpus; a seeded 2,000-repository sample drives the scan. The fetcher shallow-clones source and records the commit it received. It never installs or runs the server. A fresh crawl can change the corpus, so the portfolio keeps this dated snapshot separate from the live index.

PARSE ONCE, SHARE THE TREE

I parse TypeScript and JavaScript source with tree-sitter, then reuse the syntax tree across the code rules. Shell and path rules trace input toward sinks; the other checks inspect characters, tool descriptions and access scope. Python currently receives the hidden-character check only. The tree-sitter dependency floor is 0.25 because query execution uses QueryCursor; allowing older versions would resolve an incompatible API.

UNICODE-CONCEAL - HIDDEN CHARACTERS

Finds concealed Unicode characters in source or metadata. It does not need a parser, so this check reaches languages that the deeper rules do not yet support.

SHELL-EXEC-UNSAFE - INPUT TO SHELL

Flags parameter-derived input reaching a command interpreter. The question is whether input can reach the sink, not simply whether a repository imports a shell API.

PATH-TRAVERSAL - INPUT TO FILES

Looks for input from directly registered tool, resource or prompt handlers reaching filesystem operations without a recognized containment check. Recognizing that handler boundary cuts noise from programs reading their own configuration, but misses some aliases and helper calls.

TOOL-DESC-INJECTION - INSTRUCTIONS IN METADATA

Flags instruction-like text smuggled into tool descriptions. A structural filter separates suspicious instructions from ordinary descriptions before optional semantic judgment; a flag still does not establish exploitability.

SCOPE-OVERBROAD - EXCESS REACH

Checks access wider than a local tool needs, including network origin and filesystem reach. Deployment context matters: a wildcard origin on a public read-only service is not automatically excess access. The rule is deliberately narrower than a search for suspicious configuration keys.

REACHABLE CODE, NOT ALARM VOLUME

Tests, examples, build scripts and other non-runtime paths are excluded. On a 150-server trial, that removed 9 of 19 shell-execution findings: a flaw in a test harness is not necessarily reachable by an assistant. The cost is a false negative when a real server entry point lives under an excluded path such as scripts/. I state that trade-off instead of treating fewer findings as measured accuracy.

DISCLOSURE ENFORCED AT THE WRITE BOUNDARY

Every finding passes the publication gate before SARIF/JSONL output. High and critical findings stay withheld without a recorded notice; each finding owns its own disclosure window. Missing records and future timestamps cannot open it. Maintainer notification is not implemented in the reviewed revision, so nothing above medium has published. Aggregate withheld counts remain visible without exposing the affected source.

ONE WRITER FOR THE NIGHTLY SITE

The nightly scan is the only producer of the public site. A previous workflow redeployed older committed data over fresh results, making public counts flip depending on which writer ran last. I removed that competing writer and queue runs rather than cancel them. The trade-off is that a scanner or rendering change must wait for a complete scan before publication.

MODEL ACCESS STAYS OPTIONAL

The triage adjudicators, content-hash cache and spend guard are built and tested in isolation, but the pipeline does not call them. The model SDK is an optional extra, so deterministic scanning works without model credentials. Wiring triage before human labels exist would spend money on judgments whose quality nobody can independently score.

TESTS ARE NOT AN ACCURACY BENCHMARK

The reviewed repository reports 782 passing tests with ruff and mypy checks. Those protect contracts, parsing and publication behavior; they do not measure precision on real findings. Human labels are still missing. The next work described in the repository is to obtain those labels, score the judgments, wire useful triage and notify maintainers only about verified findings.

Alternatives considered

Execute servers in a sandbox

Trade-off: It would add runtime coverage, but also execute the untrusted code this scanner is meant to inspect. I accepted missed dynamic behavior in exchange for a source-only boundary.

Publish every finding immediately

Trade-off: It maximizes visible output while exposing maintainers to unverified high-severity claims. Withholding costs immediate visibility, but missing notice data must never become permission to publish.

Build every language’s deep rules first

Trade-off: Each language needs its own taint queries and fixtures. Covering everything first would delay the human-labelled measurement; I shipped stated TypeScript/JavaScript coverage and keep Python’s narrower checks explicit.

Keep both publication workflows

Trade-off: Ordering two writers would preserve the race between fresh scans and stale committed data. A single writer removes it, at the cost of waiting for a full scan to publish changes.

What the results show

Corpus, not scan coverage

23,157 repositories

Public trend record at 2026-10-04T09:46:20Z. This is the eligible corpus, not the number inspected.

Dated scan

1,673 servers · 916 findings

2026-10-04 aggregate, from a seeded 2,000-repository sample. Findings are source-pattern candidates, not confirmed vulnerabilities.

Publication boundary

171 public · 745 withheld

The same 2026-10-04 record. Serious findings remain withheld pending verified disclosure; a published count is not a precision score.

Repository-reported checks

782 tests

Status reported by the reviewed repository, with ruff/mypy clean. These checks protect implementation contracts; they do not establish scanner accuracy.

Limitations

  • Human-labelled precision is unmeasured. The golden-set machinery exists; an independently labelled answer key does not.
  • Four deeper rules cover TypeScript/JavaScript; Python currently receives only the hidden-character check. These are not ecosystem-wide accuracy figures.
  • Static analysis cannot observe runtime behavior. Excluded paths, helper calls and coarse containment checks can hide real problems.
  • Triage is not wired into scans, maintainer notification is not built, and no accuracy threshold gates project CI yet.
  • The early pilot’s labels were model-made and not retained. I do not use that pilot as a reproducible accuracy claim.
  • The live index changes nightly. Commands reproduce the method against the current registry, not the portfolio’s historical counts.

Reproduce it yourself

  • From a local checkout of the linked repository, create the environment: python3.12 -m venv .venv
  • Activate it: source .venv/bin/activate
  • Install the project and checks: pip install -e ".[dev]"
  • Build the current registry index: mcp-observatory crawl
  • Use the nightly sample size: mcp-observatory scan --index .cache/server_index.jsonl --sample 2000
  • Run the local checks: pytest -v; ruff check .; mypy
  • Inspect findings.jsonl, findings.sarif and summary.json under data/. Withheld source stays outside public output.

Sources checked on 2026-10-06 and tracked in CLAIMS.md. Nightly figures are dated snapshots; experiment results retain their original labels and stated limits.