Search guides, workflows, and reference pages.

Docs/mode economics

Mode economics — what does each mode cost to run?

Indicative, not a quote. The numbers on this page describe the measured file sizes, bounded replay samples, and unvalidated planning estimates. Token prices vary by provider, model, date, and discount tier. Always multiply by your own provider’s current rate — or by zero if you are running local inference.

This page exists because MISSION.md § Affordability commits to documenting mode economics honestly: a maintainer evaluating adoption should be able to make an informed decision, not discover the cost after the fact. The same data informs the long-term capacity planning for an ASF-hosted inference endpoint (see Long-term: the ASF inference endpoint).


How to read this page

What “tokens” means here

One token ≈ 0.75 words in English prose, or roughly one character in structured code or JSON. Token counts on this page use K for a thousand tokens and M for a million (so 30K = 30,000 and 1.5M = 1,500,000). Practical anchors:

Content Approximate token count
Typical bug-report body (400 words) ~530 tokens
Small PR diff (50 lines changed) ~800 tokens
Medium PR diff (300 lines changed) ~5K tokens
Large PR diff (1,500 lines changed) ~25K tokens
Mail thread, 10 messages ~3K–8K tokens

Full skill-file sizes are measured separately in the generated table below. They describe the file loaded when that skill is invoked, not the full session or the always-on plugin description cost. Referenced documents and tool output add further context. Other token ranges on this page are planning estimates; they are not measured runtime percentiles. The separately labeled replay benchmark below provides sample percentiles for its own bounded workloads. Different model tokenizers can produce different counts; no universal conversion percentage is assumed.

Measured skill-file tokens

Restamp with uv run --project tools/skill-token-count skill-token-count --write. The measurement tool documents the scope and normalization. Where the per-skill figures live and how to print them is at the end of this section.

Every non-setup skill’s figure below includes the shared reconciliation pre-flight check, and is smaller than it was before that check existed.

The check first grew the shared pre-flight block from 1,679 to 3,271 tokens — +1,608 on each of the 65 skills carrying it, +49.0% on the smallest. Three changes reversed that, each removing a layer rather than adding one. The block was split into a decision path and a cold sidecar. Then the deterministic half — read a lock, order two versions, compare two hashes, subtract two dates, decide whether a proposal was already shown — moved out of prose entirely into tools/setup-preflight, which the block runs as one command. Then the rules prose moved there too, and the command now emits the sections its own findings name, so there is no second file to read and no per-skill copy of one.

The block is 585 tokens, against 1,679 before the check existed and 3,271 at its peak. Each of the 65 skills is 1,075–1,081 tokens cheaper than on main while carrying the whole check: ci-runner-audit 3,281 → 2,203 (−32.9%), and the sidecar that briefly cost 2,516 tokens × 65 copies in the repository is gone.

The rules are 2,057 tokens across eight sections, held once in the tool. A run that needs one pays for one — typically 120 to 580 tokens — and the ordinary {"verdict": "ok"} pays for none.

What is left in the block is the part a model is for: run the command, stay silent on ok, follow the rules a finding carries, and never run /magpie-setup adopt unattended. What is left in neither is the arithmetic, which is now tested rather than graded.

Each skill’s count lives in its own SKILL.md, as a generated measured_tokens: frontmatter line, the same way surface_hash: does. It is measured with tools/skill-token-count (pinned tiktoken, cl100k_base): the full UTF-8 file, frontmatter and comments included, excluding the measured_tokens: line itself. The website renders the per-skill table from those stamps at build time; to print it locally:

uv run --project tools/skill-token-count skill-token-count --table

The table is not committed here. A single generated table covering every skill made any two skill PRs conflict on its shared lines; a per-skill stamp only conflicts when two PRs edit the same skill.

Tokenizer: cl100k_base (pinned tiktoken). Each figure is the skill’s own measured_tokens: stamp: the full file, frontmatter and comments included, excluding the stamp line itself.

Skill file Measured tokens
audit-finding-fix 6,091
ci-runner-audit 2,201
committer-onboarding 7,915
contributor-activity-sweep 3,398
contributor-calibrate 3,011
contributor-candidate-screen 2,997
contributor-identity-map 3,253
contributor-nomination 5,610
contributor-sentiment 4,720
contributor-to-committer 5,741
dependency-audit 3,110
dependency-license-audit 5,244
flaky-test-triage 3,069
good-first-issue-author 3,609
good-first-issue-sweep 4,123
issue-backlog-stats 4,729
issue-deduplicate 4,463
issue-fix-workflow 4,795
issue-reassess 4,911
issue-reassess-stats 2,922
issue-reproducer 5,043
issue-stale-sweep 4,910
issue-triage 4,973
license-compliance-audit 4,631
list-skills 2,286
mentoring-welcome 3,222
newcomer-issue-explainer 3,491
onboarding-concierge 3,374
optimize-skill 3,209
pairing-multi-agent-review 3,686
pairing-self-review 3,437
pr-management-code-review 4,925
pr-management-mentor 2,942
pr-management-quick-merge 4,717
pr-management-stats 3,434
pr-management-triage 5,005
pr-stale-sweep 4,932
pre-first-pr-check 3,388
release-announce-draft 6,907
release-archive-sweep 4,447
release-audit-report 6,629
release-keys-sync 4,799
release-prepare 13,872
release-promote 6,883
release-rc-cut 11,783
release-verify-rc 10,725
release-vote-draft 6,661
release-vote-tally 5,535
report-framework-issue 4,517
reviewer-routing 4,867
security-cve-allocate 6,485
security-issue-deduplicate 5,411
security-issue-fix 5,912
security-issue-import 9,243
security-issue-import-from-md 6,799
security-issue-import-from-pr 8,059
security-issue-import-from-scan 5,456
security-issue-import-via-forwarder 5,639
security-issue-invalidate 6,974
security-issue-sync 6,156
security-issue-triage 6,552
security-model-prepare 4,526
security-model-update 4,652
security-model-verify 5,462
security-tracker-stats-dashboard 3,546
setup 4,530
setup-isolated-setup-doctor 6,034
setup-isolated-setup-install 5,524
setup-isolated-setup-update 5,114
setup-isolated-setup-verify 4,951
setup-override-upstream 4,519
setup-privacy-llm 2,051
setup-shared-config-sync 3,700
setup-status 2,318
setup-upstream-fix 5,215
skill-reconciler 4,363
workflow-security-audit 3,554
write-skill 2,456

Measured runtime replay sample

On 2026-09-11, 30 actual Claude Code invocations ran synthetic tasks through five skills across four modes to their final draft/report boundary: three scenarios per skill, two repetitions each, using claude-haiku-4-5-20251001 and Claude Code 2.1.268. Full entrypoints and sibling Markdown references were supplied, along with captured configuration and tool observations. Live tools and posting were disabled.

Mode and measured skill Runs p50 total tokens p90 total tokens
Triage: issue-triage 6 34,825 35,055
Mentoring: pr-management-mentor 6 28,907 29,592.5
Mentoring: good-first-issue-author 6 26,567 26,653
Drafting: issue-fix-workflow 6 28,714 29,406.5
Pairing: pairing-self-review 6 22,975 24,475

Totals use CLI-reported per-model usage, including uncached input, cache creation, cache reads, output, and auxiliary CLI calls. Thinking is included in output, not added twice. These are token-traffic counts, not equivalent billing units. p50/p90 use inclusive linear interpolation. Six runs across three synthetic scenarios are an exploratory sample, not typical costs for an entire mode.

Concrete first-attempt examples:

  • Triage: classify an empty-input bug and draft a proposal: 34,962 tokens.
  • Mentoring: draft regression-test guidance for a newcomer: 29,207 tokens.
  • Mentoring, issue authoring: draft an issue for a missing test: 26,131 tokens.
  • Drafting: draft a one-file empty-input fix and regression test: 28,600 tokens.
  • Pairing: review one file whose empty-input guard was removed: 22,406 tokens.

The measurement report provides input/cache/output breakdowns and methodology. The records and final responses and synthetic corpus provide the first 24 calls. The Drafting report, records, and corpus cover six further calls on an empty-input fix, a weighted-mean denominator fix, and a two-module summary fix. All 30 CLI calls produced artifacts; completion is not a correctness grade. One two-file review miscounted deleted lines. One Drafting response asserted that tests were green inside its proposed PR text although no tests ran. These outputs need human review and are not verified fixes.

The initial sample mislabeled good-first-issue-author as Drafting. It is Mentoring; metadata has been corrected without changing prompts, usage, or responses. Legacy drafting-* case IDs in that sample refer to issue authoring. The collector now checks each case against the skill’s declared mode before launching calls. Percentiles are grouped by skill, not pooled across a mode.

These observations do not measure live GitHub/tool calls, follow-up discussions, patch application and test execution, or multi-agent pipelines. They must not be substituted for the broader planning ranges below or generalized to other models.

Security draft pre-flight also loads the shared CC-resolution rule from tools/mail-source/contract.md: approximately 600 additional tokens once per run, estimated from the rule’s prose and configuration identifiers. It reuses already-loaded project/organization configuration and adds no mail or tracker calls, so the per-mode ranges below remain unchanged.

Model classes

Skills are written against a capability contract, not a vendor. Three capability classes cover the realistic range for these workflows:

Class Parameter scale Characteristics
Small ~7B–13B equivalent Fast and cheap. Good at extraction, classification, and short structured drafts. Struggles on long-chain reasoning, large contexts, and novel patterns.
Mid-tier ~70B equivalent Balanced quality and cost. Handles the full skill catalogue well. Recommended starting point for new adopters.
Large Frontier reasoning Highest capability and highest cost. Use where mid-tier recall or reasoning falls short — complex security analysis, multi-step code fix drafting, detecting novel vulnerability patterns.

Local models (Ollama, vLLM, llama.cpp) map onto Small or Mid-tier by capability; they incur hardware cost rather than per-token billing. See Local and self-hosted inference.


Per-mode token shape

Planning estimates are unvalidated hypotheses, not observed bounds or averages. The replay p50 column refers only to the exact skill and synthetic workload above, using Claude Code’s total token traffic including cache and auxiliary calls. Not measured means there is no corresponding session evidence here. It does not mean zero, and another skill’s p50 must not be substituted.

For example, issue-triage has replay p50 34,825, above the old 4K-15K planning estimate; pr-management-mentor also exceeds its estimate. Loading full entrypoints and sibling references plus CLI overhead differs from the unspecified protocol behind those estimates. We cannot attribute the discrepancy to one cause or validate the old bounds from this sample. Use a workload-matched pilot for budgeting; the old estimates are retained for context, not as a cap.

Triage

Most Agentic Triage skills are read-bounded: the expensive part is loading context (PR diff, report body, existing issue sample), not generating output. Every output is a short proposal — a label suggestion, a routing recommendation, a classification with rationale — so output tokens are low relative to input.

Skill Typical invocation Planning estimate Primary cost driver Replay p50 tokens
pr-management-triage Single PR triage pass 5K–30K PR diff size and comment count Not measured
pr-management-stats Weekly queue report 10K–50K Number of open PRs read Not measured
pr-management-code-review Single PR deep review 15K–80K Diff size; code-heavy PRs are expensive Not measured
issue-triage Single issue classification 4K–15K Issue body length + similar-issue cross-check sample 34,825
issue-reassess Pool-level sweep (10 issues) 30K–120K Pool size; batch cost scales linearly Not measured
security-issue-import Single inbound report 8K–25K Report length + known-dup cross-check Not measured
security-issue-import-from-pr Single security PR import 10K–30K PR diff + associated discussion Not measured
security-issue-import-from-md Batch import (5 findings) 15K–60K Number of findings × finding length Not measured
security-issue-deduplicate Two-tracker merge 10K–30K Tracker age and mail-thread depth Not measured
security-issue-invalidate Single invalid close 8K–20K Report length + reply draft Not measured
security-issue-sync Full tracker reconciliation 20K–100K Tracker age, mail-thread depth, linked PRs Not measured
security-cve-allocate CVE allocation workflow 5K–12K Mostly procedural; low variance Not measured
security-model-verify One repository, both checks 10K–40K Reads the chain plus the whole model document; a shared model is read once for the repository set Not measured
contributor-identity-map One contributor’s channel identities 5K–20K Number of reachable sources and name-match candidates Not measured
contributor-activity-sweep Single-contributor activity card 10K–40K Activity volume in the configured window Not measured
contributor-sentiment Full sentiment gate report 20K–80K Number of threads and signals sampled Not measured
contributor-nomination Nomination-readiness brief 20K–70K Contributor activity breadth read, plus the conversations on up to 50 authored items, 50 comment threads and 20 reviews for the automated-contribution discount, and any project expectation documents Not measured

Illustrative planning assumption, not a measured average: assuming 10K-30K tokens per item and 50 items per week gives 500K-1.5M tokens/week. The measured replay above exceeds that per-item assumption; validate it for your workload.

Mentoring

Agentic Mentoring is conversational and per-reply: the agent reads thread context, project conventions, and contributor history, then produces a single targeted response. Cost per reply is moderate; total weekly cost depends on contributor volume.

Skill Typical invocation Planning estimate Notes Replay p50 tokens
pr-management-mentor Single threaded reply 6K–20K Estimated; skill experimental 28,907
good-first-issue-author One candidate → one issue draft 6K–18K Estimated; reads one candidate + named source files, no full-thread history; skill experimental 26,567
newcomer-issue-explainer One issue → one beginner explanation draft 4K–12K Estimated; reads one issue body + a small set of named source files; read-only; skill experimental Not measured
mentoring-welcome One first-time contributor → one welcome draft 4K–12K Estimated; reads the triggering thread + contributing-guide pointers, no full-thread history; skill experimental Not measured
onboarding-concierge One newcomer question → one grounded answer draft 4K–12K Estimated; reads CONTRIBUTING.md + the relevant doc excerpt; read-only; skill experimental Not measured
contributor-to-committer Single contributor readiness brief 20K–70K Estimated; reads the contributor’s activity history against the adopter’s thresholds, plus the conversations on up to 50 authored items, 50 comment threads and 20 reviews for the automated-contribution discount, and any project expectation documents; read-only; skill experimental Not measured
contributor-calibrate One calibration run (~100 past nominations) 150K–500K Estimated; reads the private list’s nomination threads (rows only are kept), then runs contributor-metrics per nominee and window with up to 10 pushback confirmations each; read-only until the confirmed config diff; skill experimental Not measured
contributor-candidate-screen One screening run (pool of ~100, shortlist of ~15) 200K–800K Estimated; count-only pre-filter over the pool, then contributor-metrics and community signals per survivor and shortlisted candidate, plus two or three paragraphs each; writes only after confirmation; skill experimental Not measured
good-first-issue-sweep Backlog sweep (10 issues) 20K–80K Estimated; scales linearly with the number of issues scored; skill experimental Not measured

Unvalidated planning assumption for Agentic Mentoring: budget 10K–20K tokens per contributor interaction. A project with 20 active contributors each receiving 3 agent replies per week: roughly 600K–1.2M tokens/week.

Drafting

The most variable mode. Short reporter replies are inexpensive; agent-drafted code fixes are expensive because the agent reads relevant source files in addition to the issue or report.

Skill Typical invocation Planning estimate Notes Replay p50 tokens
security-issue-fix — reporter reply Single reply draft 10K–35K Reads report + canned responses + prior thread Not measured
security-issue-fix — code fix Agent-drafted fix + PR 30K–150K Adds source files; wide variance Not measured
issue-fix-workflow Issue fix + PR 25K–120K Bounded by what the skill reads from the codebase 28,714 (patch draft only)
security-model-update One update cycle 40K–200K Dominated by the corpus: closed trackers with discussion, the reporter threads, and the model itself. Scales with the window, not the diff Not measured
security-model-prepare First model for one project 150K–600K+ The deep surface pass over in-scope entry points is the cost, and it is meant to be — a model written from the README alone is a summary of marketing copy. Budget it as a project, not an invocation Not measured

Unvalidated planning assumptions for Agentic Drafting: 15K-25K tokens for reporter replies and 50K-100K tokens for code-producing invocations depending on codebase scope. Limiting the skill to the relevant source files is the single biggest lever on Agentic Drafting cost.

security-model-prepare is the outlier in this table and is best budgeted separately: it is a one-off per project, its cost is front-loaded into a code-reading pass whose whole purpose is to be thorough, and what it produces is amortised across every later triage decision the model routes. The recurring cost is security-model-update, which is bounded by the window it is given.

Pairing

Agentic Pairing runs in the developer’s own development cycle, not on project infrastructure — cost is per-developer-session. Multi-agent pipelines multiply the per-pass cost by the number of review agents. Whether a project reimburses contributors is a project policy decision. The following ranges are estimates, not measured session costs.

Skill Typical invocation Planning estimate Notes Replay p50 tokens
pairing-self-review Pre-flight review of a local diff 10K–60K Estimated; skill experimental. Scales with diff size plus conventions, dependency, and release-policy doc length. 22,975
pairing-multi-agent-review Full three-pass review 30K–200K Estimated; skill experimental. 3–4 × single-pass cost. Parallelism reduces latency, not billing. Not measured
pre-first-pr-check Newcomer pre-flight checklist on a local branch 5K–20K Estimated; skill experimental. Read-only; scales with diff size and convention docs read. Not measured

Unvalidated planning assumptions for Agentic Pairing: 15K-35K tokens for a medium-PR self-review and 45K-90K for a three-agent review. The replay measures only small self-reviews; the pipeline remains unmeasured.

Meta

The Meta mode is mostly framework machinery — setup, utilities, dashboards — whose cost is per machine or per repo rather than per maintainership item, so it is not budgeted per invocation here (see docs/modes.md § Meta). The exception with a recurring per-item shape is committer-onboarding, which runs once per new committer or PMC member:

Skill Typical invocation Planning estimate Primary cost driver Replay p50 tokens
committer-onboarding One post-vote onboarding walkthrough 10K–30K Mostly procedural; podling vs TLP path length Not measured

Modes not yet covered

Agentic Autonomous, formerly Auto-merge: off, not implemented; no measured token cost.

Model class and mode cost shape

The table below describes the quality/cost trade-off per mode, not a hard recommendation. “Viable” means acceptable recall on typical cases; “Recommended” means the sweet spot between quality and cost; “Large class” means quality requirements that mid-tier models often miss.

Mode Small class Mid-tier class Large class
Agentic Triage — classification / routing Viable for most cases Recommended default Rarely needed
Agentic Triage — security import (novel patterns) Miss rate is higher Recommended default For subtle or novel reports
Agentic Mentoring Acceptable on simple threads Recommended default Not typical
Agentic Drafting — reporter reply Acceptable Recommended default Rarely needed
Agentic Drafting — code fix Often insufficient Recommended default Complex bugs or large refactors
Agentic Pairing — self-review Limited recall on conventions Recommended default Anchor pass in multi-agent pipelines

Price comparison: this page does not maintain a dated comparison of named models, so it does not claim a current price multiplier between classes. For a concrete budget, apply the selected provider’s current input, output, and cached-input rates to the corresponding measured token quantities. The class labels alone do not establish an invocation’s price.


Local and self-hosted inference

Running a model locally (Ollama, vLLM, llama.cpp) shifts cost from per-token billing to hardware:

Inference path Per-token cost Typical hardware cost Notes
Consumer GPU, Small-class quantised model $0 ~$0.10–0.50/hr (capex amortised over ~3 yr lifespan × moderate utilisation) Viable for Agentic Triage and short Agentic Mentoring/Agentic Drafting
Cloud spot GPU, Mid-tier model $0 ~$1–4/hr depending on GPU class Viable for all modes; latency is higher than hosted APIs
CPU-only, quantised Small model $0 Near-zero Very slow; not recommended for interactive Agentic Pairing

Local inference is also the simplest privacy answer for most skills: data never leaves the machine, and no third-party data-processing agreement is needed. The framework’s vendor neutrality means local paths use identical skill code to hosted paths.


Reducing costs

  1. Match model class to task. Agentic Triage classification and short Agentic Mentoring replies do not need a frontier model. Reserve Large-class for novel-pattern security analysis and complex multi-file code fixes.

  2. Scope code reads. The biggest driver of Agentic Drafting cost is how many source files the agent loads. Small, well-named files help the skill read only what is relevant.

  3. Cache skill context. Most agent CLIs support prompt-level caching. The skill file (size varies by skill class; see What “tokens” means here) and stable project configuration files are ideal cache candidates — the first invocation pays; subsequent invocations are cheap on the cached portion. Note: most provider caches have a short TTL (Anthropic prompt cache: 5 min default, 1 h extended at higher write cost), so bursty same-session workloads benefit most; periodic triage runs spaced hours apart will typically miss the cache.

  4. Batch triage. issue-reassess and pr-management-stats amortise context load across a pool. Running them weekly rather than per-event reduces overall token volume compared with individual calls.

  5. Run locally for development. When authoring or testing a new skill override, use a local model. Save the hosted model for production invocations.


Long-term: the ASF inference endpoint

MISSION.md § Affordability names an ASF-hosted inference endpoint as a long-term roadmap item: a community-affordable, foundation-governed, audit-logged inference layer any open-source maintainer — ASF or otherwise — can use without paying a vendor or accepting a vendor’s gift.

Part of that now exists. LLMAO (llm.apache.org) went live in September 2026 as the Foundation’s sanctioned-inference gateway: a committer authenticates with a personal access token, and it serves three self-hosted models. The gateway’s defaults, model list and known limitations are recorded in organizations/ASF/organization.md.

Model Context Reasoning on by default
gemma4-26b (recommended) 131,072 No
qwen3.8-27b 131,072 Yes
qwen3-8b 40,960 Yes

Two caveats that matter for this page specifically:

  • It is a pilot, and its privacy class is project-internal. LLMAO serves from rented third-party GPU hardware rather than ASF-operated infra, and pilot traffic must be treated as visible to gateway admins. It is carved out of the *.apache.org default approval, so it is not approved for <private-list> or <security-list> content. See tools/privacy-llm/models.md.
  • Tool use over the Anthropic-compatible path is currently broken upstream. LiteLLM routes it to vLLM’s /v1/responses with a tool_choice shape vLLM rejects. Every Magpie skill is tool-driven, so the replay benchmark above cannot run against the gateway until that lands. Plain conversation is unaffected.

Planned: measuring the modes against LLMAO

No LLMAO figures appear on this page yet, and none should be inferred from the throughput numbers the gateway publishes — those come from synthetic load, not from skill workloads.

The intended run reuses the harness that produced the replay sample above, so the results are comparable rather than a separate methodology: the same corpus and scenarios, --model pointed at each of the three gateway models, with ANTHROPIC_BASE_URL and a committer PAT in the environment. What it would add to this page is the thing the model-class table currently asserts from capability reasoning rather than measurement — whether a ~26B self-hosted model actually carries the mid-tier workloads, and which modes degrade first when it does not.

It is gated on the tool-use limitation above; that is the first thing to re-test, because it decides whether the measurement is possible at all. Tracked in issue 1260.

The file counts and bounded replays on this page are initial evidence for the capacity planning and cost models a foundation-governed endpoint will need. The planning estimates are not validated capacity requirements. As pilot adopters accumulate real usage data, this page will be updated with observed ranges rather than theoretical estimates, so the endpoint sizing argument rests on evidence.


Cross-references

Suggest a change