tools/dev/
Capability: substrate:framework-dev
Harness: agnostic
Framework dev-loop helpers (placeholder check, agent pre-commit hook). Invoked by prek and CI; not consumed by any skill directly. See the individual scripts in this directory for usage.
The shared dev toolchain
tools/dev is also the workspace’s toolchain project, magpie-dev. It
declares ruff, mypy, and pytest as its dependencies, and every other workspace
member names magpie-dev in its own [dependency-groups] dev instead of
repeating the pins. Bump a version here and the whole workspace moves together.
Two deliberate exceptions: tools/vetted-ops and tools/adversarial-review
declare no dev group. Each ships as a plugin and runs from outside the
workspace, where magpie-dev cannot resolve; their tests get the toolchain
from the root dev group instead.
Each member’s environment stays self-contained — the checks still run
uv run --directory <member> --project . python -m <tool>, so no member depends
on tools leaking in from the root environment. Only the declaration is shared.
Why it changed: the pins used to be repeated in every member with an instruction to keep them in lockstep. They had drifted into three different mypy floors, two pytest floors, and two ruff floors, one member had no dev group at all, and nothing detected any of it — the duplication was the bug, and the instruction to keep it consistent was the workaround.
The project builds as a metadata-only wheel: it ships no importable module, because the scripts are hyphenated and invoked by path, but it has to be installable for other members to depend on it.
The scripts
| Script | What it does |
|---|---|
check-companion-skills.py |
Generates the Works well with block in each family README and the companion-skills page from companion-skills.json — third-party skill packages that pair with a family, none of them a dependency. Generated because the advice is per family and per harness: Superpowers installs on six agents with six different commands, Claude Security on one. As prose across ten READMEs that is forty-odd facts to keep straight, and the first to rot tells a Codex user to run a Claude Code command. Validates the registry shape, that every claimed family exists, that every named harness is one Magpie documents an install for, and refuses an entry whose why cannot say what it adds to a named family — an entry that cannot is an advert. Each harness entry carries install and marketplace separately: adding somebody else’s catalogue is a trust decision rather than an implementation detail of an install command, so the install flow can name whose it is and ask about it on its own. marketplace is null where there is none to add — the package is already in a catalogue the agent has, or that harness installs from a URL — and naming one for a harness with no marketplace concept fails. Also the place the vendor-neutrality framing is enforced by construction: every entry prints whose tool it is and which agents can run it. |
check-skill-config.py |
Generates each family README’s Before the first run config table from the skills’ requires_config: frontmatter, and guards it. Which <project-config>/ files a skill reads, and whether it can work without one, used to live in prose (only 8 of 74 skills carry a config section, under three headings, marking optionality however the author felt like) and in a hand-maintained table per family — which had drifted as far as an unguarded table does: security listed none of the thirteen files its skills read, repo-health none of three, release-management one of five. The required set is now declared once per skill in frontmatter; the optional set is derived (referenced minus required), so the larger half never needs maintaining. Also checks that a declared file is actually read by that skill, and that every file any skill reads has a template in projects/_template/ — a skill must not ask an adopter for config the framework cannot scaffold, which is how three such files were found. Descriptions come from the adopter scaffold’s own index, falling back to the template’s title, so this introduces no third copy of “what this file is for”. --fix rewrites the blocks. |
check-doc-sync.py |
Guards the documentation claims that track the tree and rot silently: spec-index completeness (every tools/spec-loop/specs/*.md listed in both overview.md and README.md), the per-family skill counts in the root README.md, the per-mode counts in docs/modes.md’s Modes at a glance table, the bare catalogue totals in docs/setup/marketplace.md, the per-family plugin counts in the marketplace tables of docs/setup/marketplace.md and docs/quick-start.md (bare integers in a table column, which the other count checks do not match), the “one plugin, N skills” claim in each family README’s Install & first runs section (keyed on the install command, not the directory name), that no doc invokes a family skill by a name repeating its family (/magpie-<family>:<family>-…, which the plugin does not advertise), that the published always-on token figures match estimate-skill-tokens.py, that a doc showing the portable single-token form (/magpie-<skill>) says which install it means — /magpie-setup is exempt as the name of the install mechanism, and filesystem paths are not invocations — and that every script here is named in this file. Eleventh check: the declared eval-case counts — each tools/skill-evals/evals/<family>/README.md headline total and per-suite row, and each family’s line in tools/skill-evals/README.md — against the fixtures/case-*/ directories the runner actually walks, plus that every eval family has an index entry at all. --fix rewrites those numbers and generates a missing entry; the suite-name list in the parenthetical is left alone, being prose rather than a count. Twelfth check: the one-time marketplace add — /plugin marketplace add apache/magpie and the Codex, Gemini and apm equivalents — appears only in docs/setup/marketplace-install.md and docs/setup/marketplace.md; docs/designs/ is exempt, because a design records what was decided. The per-family /plugin install magpie-<family>@apache-magpie line is not guarded: it differs per page and is the page’s subject. |
render-wizard.py |
Generates every animated SVG the docs show. Top level: the quick start’s hero (install.svg — three install commands and a skill answering, the shortest true story), one per walkthrough step (step-install, step-isolation, step-use, step-adopt), and magpie-setup.svg, the whole first run including the secure-agent setup. Per family under assets/quickstart/wizard/: /magpie-setup config playing through. The static screenshots show a skill’s output; they cannot show a conversation, and a still frame of a wizard is a wizard with the interesting part removed. Every frame is derived, not written: which files that family’s wizard would create comes from its skills’ requires_config: frontmatter, and the one-line description of each from the adopter scaffold’s index — the same two sources check-skill-config.py reads, so a family that gains a required file gains a frame with nobody editing a transcript. Animation is SMIL (<animate> on opacity, one element per line sharing one duration so the sequence loops as a unit): plain XML, deterministic, and no Node at all. A renderer that does not animate shows the first frame, which is the command about to be typed. Illustrative rather than a recording, and the embed says so. --check fails on drift; check-quickstart-recording.py calls it. It replaced record-svg.sh, whose last target this was: the committed recording opened with the marketplace install — a prerequisite with its own page since — and re-cutting it needed a terminal, a scratch project and a human, which is why it stayed wrong. Nothing in this repository is captured any more. |
render-screenshot.sh |
Renders an authored terminal transcript (assets/quickstart/families/<family>/<name>.txt) to the static SVG beside it. The family screenshots are written, not captured: a capture needs a terminal, a scratch project and a human, and needs all three again whenever any output moves — which is what left nine family recordings showing the same thing for months. What a capture gives that authoring does not is a guarantee that the picture matches the program, and this buys back the half it can: rendering is deterministic — same .txt, same bytes — so check-quickstart-recording.py can prove every committed .svg still matches its source. Colour comes from the line itself (a > prompt, a ✓/⚠/✗ status glyph, an ALL-CAPS section label), so a transcript stays something you read as a terminal rather than as markup. A leading block of # comments is stripped before drawing, so each transcript carries the ASF licence header like any other authored file and Apache RAT needs no .rat-excludes entry for it — verified by the rendered SVGs staying byte-identical when the headers were added. Shares one palette with render-wizard.py, so the static and animated sets read as one terminal. --all renders every transcript; --check fails on a stale or unpaired file. No Node, on a bare clone and in CI. |
check-quickstart-recording.py |
Validates the one recording and the authored screenshots against what the repo actually ships. The recording (assets/quickstart/magpie-setup.svg) must exist and be embedded by the quick start and by docs/setup/README.md. The screenshots (assets/quickstart/families/<family>/<skill>.{txt,svg}): every live family has a directory with at least one transcript, every directory belongs to a family, every .txt has an .svg and vice versa, every name is a skill that family actually ships, and every .svg is embedded by its family README. The load-bearing one is regeneration staleness — it shells out to render-screenshot.sh --check, so an .svg edited without its .txt (or a .txt fixed without a re-render) fails the build. That is the guard a hand-authored image cannot have, and the reason the renderer is byte-deterministic. All SVGs: parses as XML with an <svg> root, carries the Apache licence header, under the 1536 KB cap. Also guards three retirements — the fourteen PNG stills, the nine *-first-run.svg recordings that showed setup’s arc rather than the family’s, and the assets/examples/ set the authored screenshots replaced — so a doc referencing any of them, or a revived directory, fails. docs/designs/ is exempt: a design records what was replaced. There is no placeholder state any more; a missing file is an error. It also covers the first-run walkthrough (assets/quickstart/walkthrough/), where the constraint is different: those are numbered steps a reader follows top to bottom, so besides pairing and staleness it checks that docs/quick-start/first-run.md embeds every one in order — a page whose pictures are a step out of sequence teaches the wrong thing while every link in it still resolves. Runs as the check-quickstart-recording pre-commit hook. |
check-shared-blocks.py |
Owns every shared prose block that would otherwise be hand-copied across skills — replaces the single-block check-skill-preflight.py. The auto block (preflight) keeps that retired script’s exact behaviour byte-for-byte: generated from preflight-block.md, auto-inserted after the first body # heading of every non-setup-family skills/*/SKILL.md, removed from a skill that becomes exempt. Any number of declared blocks can be added under tools/dev/blocks/<name>.md; a target opts in by already carrying a delimited <!-- BEGIN MAGPIE BLOCK: <name> ... --> region (empty or filled), and this tool only ever fills that region — it never inserts one, and a target naming a block with no matching source is a hard error, never a silent skip. Declared targets are restricted to the skills/ tree. --fix propagates; bare invocation reports and exits non-zero on drift. |
skill-surface-hash.py |
Stamps a surface_hash: reconciliation fingerprint into every skills/*/SKILL.md, folding requires_config: (order-independent) and the skill’s structural anchors (##/### headings and **Golden rule ...** callouts, markdown-decoration-stripped) into a short sha256: digest. Deliberately excludes the shared pre-flight block and ordinary prose, so a reworded paragraph never moves the hash but a renamed step or an added config dependency does — a running skill has no other way to know whether its own surface moved since the project was last reconciled against it. Unlike check-shared-blocks.py, exempts nothing: the setup family’s own surface can drift too. --fix writes the field; runs after the shared-blocks hook and before the token-count check. |
check-duplication.py |
Fails the build on new cross-file near-duplicate prose across skills/ (recursively), tools/dev/blocks/*.md, and preflight-block.md — the gate that stops the duplication check-shared-blocks.py removes from coming back. Paragraphs over 25 words, normalised to lowercase word tokens and compared across files as sets of 9-grams, scored |A ∩ B| / min(|A|, |B|); fails above 0.50, reports (without failing) everything in [0.30, 0.50], says nothing below that. Reuses check-shared-blocks.py’s own PREFLIGHT_RE / DECLARED_RE marker regexes to blank out generated regions before scoring — blanked rather than deleted, so removing one can never fuse the paragraph before it onto the paragraph after it — and excludes YAML frontmatter and fenced code blocks the same way. A failure names both files, both line numbers, the score, and a snippet of each paragraph, and points at tools/dev/blocks/<name>.md as the remedy. |
estimate-skill-tokens.py |
Estimates each marketplace plugin’s always-on token cost — the frontmatter name + description every installed skill advertises on every turn, at ~4 chars/token — and prints it per family. --check compares the figures published in docs/setup/marketplace.md and docs/quick-start.md against the live frontmatter; check-doc-sync.py calls it, so an edited description that moves a published number fails the build. The SKILL.md body is excluded: it costs nothing until the skill is invoked. |
check-family-plugins.py |
Validates the marketplace plugins against the skills’ family: frontmatter — version parity across every ecosystem manifest, Agent Plugins 1.0 conformance, and one well-formed per-family plugin whose skills/ symlinks match the family exactly. --fix regenerates them, which is how the prek hook runs it. |
bump-dev-version.py |
Moves the .dev<YYYYMMDDHHMM> stamp on project.version in the root pyproject.toml — the single authority every manifest mirrors. Only the edit: check-family-plugins.py --fix and uv lock still follow, so the three steps read the same whether a human or CI runs them. The stamp is UTC at minute resolution and has to move for adopters to pick anything up (claude plugin update compares version strings, so a frozen suffix is a silent no-op). Refuses a version with no .dev suffix rather than stamping a release. Called by bump-dev-version.yml. |
gh-signed-commit.py |
Commits the working tree through GitHub’s createCommitOnBranch mutation instead of git commit + git push, so the commit is signed by GitHub and shows as Verified with no key material in CI. Collects changed and deleted paths from git status --porcelain -z, decomposes renames (the mutation has no rename concept), and pins expectedHeadOid so a concurrent push fails the call rather than being overwritten. Used by bump-dev-version.yml. |
check-placeholders.sh |
Fails the build on hardcoded project references in skill and tool docs, which must use <PROJECT> / <project> / <tracker> / <upstream> instead. Carries both casings and matches spaced variants. |
check-workspace-members.py |
Catches a new tools/<name>/pyproject.toml that was never added to [tool.uv.workspace] members — an omission that silently drops the tool from both the pre-commit hooks and the CI pytest matrix. Also verifies each member’s tests actually run: both surfaces key off [tool.pytest.ini_options], so a project can carry a full tests/ directory and be executed by nothing. Reports tests-without-config, config-without-tests, and neither; [tool.magpie.checks] skip = ["pytest"] is the declared exemption. |
generate-labeler-config.py |
Generates .github/labeler.yml — the path → label map the labeler.yml workflow uses to pre-apply contract:* / substrate:* labels to a new PR — from the **Capability:** line of every tools/<name>/README.md, so a new tool or a changed capability needs no hand-edit. A skill’s eval fixtures under tools/skill-evals/evals/ and the spec-loop specs do not count as touching their tool. Rewrites in place and exits 1 on change, which is how the prek hook runs it; --check only reports. |
run-workspace-check.sh |
Runs one static-check or test command across every workspace member, auto-discovering which members a given check applies to. The four workspace-* hooks call it, so adding a tool needs no edit to the pre-commit config. |
add-license-headers.py |
Stamps the SPDX licence header into Markdown files that lack one. |
agent-pre-commit.sh |
Wrapper for prek run --all-files, for agent use. An agent running pytest / ruff / mypy individually still misses the rest of the CI gate (doctoc, markdownlint, typos, the checks above); this runs what CI runs. |
Each check-doc-sync.py check was added after the drift it catches had been
found by hand. None of them break anything when wrong, which is precisely why
they need a machine rather than a reviewer: they are numbers and index entries
a human has to remember to update while thinking about something else.
Prerequisites
- Runtime: Bash + coreutils;
check-workspace-members.py,check-family-plugins.py,check-doc-sync.py,check-quickstart-recording.py,check-shared-blocks.py,skill-surface-hash.py,check-duplication.py,generate-labeler-config.py, andadd-license-headers.pyrun underpython3(standard library only).check-quickstart-recording.pyparses the recording withxml.etree, so the check needs no image library; every SVG it validates is generated byrender-screenshot.shorrender-wizard.py, neither of which needs Node. - CLIs:
uv(the workspace checks runuv run),git, andprek(orpre-commit) — these scripts wire up the framework’s hooks. - Credentials / auth: None.
- Network: Local checks;
uvmay resolve workspace dependencies from PyPI (pypi.org,files.pythonhosted.org) on first sync.