Which agent is the best subagent?

I gave six agent setups one real feature in Bruno, then graded the code against the same rubric. Fable scored highest. Codex was cheapest. GLM via Claude Code sits between them.

678910Score / 10 ↑$0$0.25$0.50$0.75$1.00$1.25$1.50Estimated subscription dollars per taskFable 5.1 lowClaude CodeOpus 5 lowClaude CodeGLM 5.3Claude CodeGLM 5.3ZCodeGPT-6 Astra lowCodex CLISonnet 5 mediumClaude Code
AnthropicZ.aiOpenAI

The results

AgentScore / 10Subscription / taskWall timeAPI-equivalent
Fable 5.1 low9.75$0.8529 min$15.85
Opus 5 low9.5$1.0729 min$18.81
GLM 5.3 via Claude Code9$0.2441 min$5.64
GLM 5.3 via ZCode8.75$0.1936 min$4.39
Codex GPT-6 Astra low7.75$0.1714 min$6.05
Sonnet 5 medium7$0.7841 min$13.57

What I would use

  • Fable for the highest score. 9.75 out of 10 for about $0.85. Opus scored 9.5 and cost $1.07. Fable was better and cheaper in this evaluation.
  • GLM for the middle ground. Claude Code scored 9 for about $0.24. ZCode averaged $0.19 per session and scored 8.75 on the shared final branch. Its average session cost was lower, with a slightly lower grade.
  • Codex for the cheapest run. 14 minutes, $0.17 and 7.75 out of 10. It had the cleanest code lane, but strict parsers and a failed live report check held it back.
  • Sonnet is off the frontier. 7 out of 10 for $0.78, tied for slowest at 41 minutes. It labelled an ARS total as USD and never passed bill answers through to sizing. It was not the most expensive.

The curve connects the cost-quality frontier: Codex Astra low, GLM via ZCode, GLM via Claude Code and Fable. It is a visual guide between observed configurations, not a prediction of intermediate results.

Where the cost comes from

Tokens are observed. Subscription dollars are estimates. Each model’s token usage is priced at its list rates, then multiplied by the monthly plan fee divided by its estimated monthly API-equivalent allowance. Automatic review calls are included at their own model’s rates.

Fable 5.1 low

Opus 4.7 automatic review

Token typeTokens$ / 1MAPI $
Fresh input21$5$0.00
Cache write, 5 min101.0K$6.25$0.63
Cached input449.2K$0.5$0.22
Output18.6K$25$0.46
API-equivalent $1.32Estimated subscription $0.08

Fable 5.1

Token typeTokens$ / 1MAPI $
Fresh input1.9K$10$0.02
Cache write, 1 hour356.3K$20$7.13
Cached input16.4M$0.25$4.09
Output65.9K$50$3.29
API-equivalent $14.53Estimated subscription $0.77

Run total: $15.85 API-equivalent, $0.85 estimated subscription. 17.4M tokens; 3.7% of a week’s allowance.

Opus 5 low

Opus 4.7 automatic review

Token typeTokens$ / 1MAPI $
Fresh input41$5$0.00
Cache write, 5 min210.3K$6.25$1.31
Cached input986.9K$0.5$0.49
Output31.1K$25$0.78
API-equivalent $2.59Estimated subscription $0.15

Opus 5

Token typeTokens$ / 1MAPI $
Fresh input262$5$0.00
Cache write, 1 hour328.4K$10$3.28
Cached input22.6M$0.5$11.28
Output66.3K$25$1.66
API-equivalent $16.22Estimated subscription $0.93

Run total: $18.81 API-equivalent, $1.07 estimated subscription. 24.2M tokens; 4.7% of a week’s allowance.

GLM 5.3 via Claude Code

GLM 5.3

Token typeTokens$ / 1MAPI $
Fresh input279.8K$1.4$0.39
Cached input18.2M$0.26$4.73
Output118.7K$4.4$0.52
API-equivalent $5.64Estimated subscription $0.24

Run total: $5.64 API-equivalent, $0.24 estimated subscription. 18.6M tokens; 1.9% of a week’s allowance.

GLM 5.3 via ZCode

GLM 5.3

Token typeTokens$ / 1MAPI $
Fresh input143.2K$1.4$0.20
Cached input15.2M$0.26$3.95
Output53.3K$4.4$0.23
API-equivalent $4.39Estimated subscription $0.19

Session average: $4.39 API-equivalent, $0.19 estimated subscription. 15.4M tokens; 1.5% of a week’s allowance.

Codex GPT-6 Astra low

GPT-6 Astra

Token typeTokens$ / 1MAPI $
Fresh input105.8K$10$1.06
Cached input4.1M$1$4.08
Output18.3K$50$0.92
API-equivalent $6.05Estimated subscription $0.17

Run total: $6.05 API-equivalent, $0.17 estimated subscription. 4.2M tokens; 0.4% of a week’s allowance.

Sonnet 5 medium

Opus 4.7 automatic review

Token typeTokens$ / 1MAPI $
Fresh input35$5$0.00
Cache write, 5 min121.2K$6.25$0.76
Cached input719.6K$0.5$0.36
Output16.3K$25$0.41
API-equivalent $1.53Estimated subscription $0.09

Sonnet 5

Token typeTokens$ / 1MAPI $
Fresh input364$2$0.00
Cache write, 1 hour471.5K$4$1.89
Cached input44.4M$0.2$8.89
Output126.6K$10$1.27
API-equivalent $12.04Estimated subscription $0.69

Run total: $13.57 API-equivalent, $0.78 estimated subscription. 45.9M tokens; 3.4% of a week’s allowance.

The subscription assumptions

PlanMonthly feeMonthly API-equivalent allowance
Claude Max · Fable$100.00$1,875
Claude Max · Opus / Sonnet$100.00$1,750
GLM Pro annual$56.00$1,300
Codex Pro$200.00$7,000

These are the subscription comparison’s existing allowance assumptions. The Claude $100 estimates are one quarter of the measured $200 plan allowances. GLM’s $1,300 estimate came from three weekly exhaustion windows. Codex uses the comparison’s central case.

The usage audit did not remeasure those allowances. These costs allocate a subscription fee; they are not per-task charges. Weekly share is the estimated subscription cost divided by the monthly fee, multiplied by 4.33.

One feature, six configurations

The task was to rebuild Bruno’s public chat bot so a price question runs the real solar estimator in conversation form. The configurations started from the same base commit in isolated worktrees. For ZCode, cost and token usage are the arithmetic mean of two sessions: $0.21 and $0.17 become $0.19 per session. The 8.75 grade and recorded wall time describe their shared final branch, not independently graded attempts.

Fable and Opus ran at low effort, Sonnet at medium, and Codex Astra at low. GLM used Claude Code’s default thinking and ZCode’s max reasoning variant. These are the configurations tested, not identical reasoning budgets.

One Fable 5.1 reviewer graded the branches with file-level evidence and re-ran the chat tests. Each rubric item scores 0 to 2; the total is divided by 2 for a maximum of 10. Self-reports were treated as claims, not evidence.

The 10-item rubric
  1. Intent gate catches all 7 openers in both languages
  2. Step machine, parsers, skips known data
  3. Reuses the wizard's geocoder, numbered pick of at most 3
  4. 'no sé' falls back to the published sample size
  5. Lead and report through the real libraries, correct money
  6. Prompt cleanup: no price dump, no early visit, 'nosotros'
  7. Interruption answered, then the pending question repeated
  8. Tests: openers, happy path, unknown consumption, interruption, ambiguity
  9. Gate green, lane respected, trailer, not merged
  10. Real end-to-end transcript and an honest report

The code under test is private, so the exact task cannot be re-run from the public repository. The redacted prompt, grading evidence, reports and extraction scripts are public. One feature evaluation does not establish a universal model ranking.

The cost correction

The first version double-counted repeated assistant-message records, priced one-hour cache writes as five-minute writes, and priced automatic review calls as the main model. This version counts each message once, uses the correct cache duration and prices each model separately.

GLM via Claude Code uses its final aggregate usage summary. ZCode averages the unique provider request totals from its two sessions. Codex uses its final cumulative turn usage, with cached input separated from fresh input. Anthropic’s list prices supply the model and cache rates.