Which agent is the best subagent?
I gave six agent setups one real feature in Bruno, then graded the code against the same rubric. Fable scored highest. Codex was cheapest. GLM via Claude Code sits between them.
The results
| Agent | Score / 10 | Subscription / task | Wall time | API-equivalent |
|---|---|---|---|---|
| Fable 5.1 low | 9.75 | $0.85 | 29 min | $15.85 |
| Opus 5 low | 9.5 | $1.07 | 29 min | $18.81 |
| GLM 5.3 via Claude Code | 9 | $0.24 | 41 min | $5.64 |
| GLM 5.3 via ZCode | 8.75 | $0.19 | 36 min | $4.39 |
| Codex GPT-6 Astra low | 7.75 | $0.17 | 14 min | $6.05 |
| Sonnet 5 medium | 7 | $0.78 | 41 min | $13.57 |
What I would use
- Fable for the highest score. 9.75 out of 10 for about $0.85. Opus scored 9.5 and cost $1.07. Fable was better and cheaper in this evaluation.
- GLM for the middle ground. Claude Code scored 9 for about $0.24. ZCode averaged $0.19 per session and scored 8.75 on the shared final branch. Its average session cost was lower, with a slightly lower grade.
- Codex for the cheapest run. 14 minutes, $0.17 and 7.75 out of 10. It had the cleanest code lane, but strict parsers and a failed live report check held it back.
- Sonnet is off the frontier. 7 out of 10 for $0.78, tied for slowest at 41 minutes. It labelled an ARS total as USD and never passed bill answers through to sizing. It was not the most expensive.
The curve connects the cost-quality frontier: Codex Astra low, GLM via ZCode, GLM via Claude Code and Fable. It is a visual guide between observed configurations, not a prediction of intermediate results.
Where the cost comes from
Tokens are observed. Subscription dollars are estimates. Each model’s token usage is priced at its list rates, then multiplied by the monthly plan fee divided by its estimated monthly API-equivalent allowance. Automatic review calls are included at their own model’s rates.
Fable 5.1 low
Opus 4.7 automatic review
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 21 | $5 | $0.00 |
| Cache write, 5 min | 101.0K | $6.25 | $0.63 |
| Cached input | 449.2K | $0.5 | $0.22 |
| Output | 18.6K | $25 | $0.46 |
Fable 5.1
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 1.9K | $10 | $0.02 |
| Cache write, 1 hour | 356.3K | $20 | $7.13 |
| Cached input | 16.4M | $0.25 | $4.09 |
| Output | 65.9K | $50 | $3.29 |
Run total: $15.85 API-equivalent, $0.85 estimated subscription. 17.4M tokens; 3.7% of a week’s allowance.
Opus 5 low
Opus 4.7 automatic review
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 41 | $5 | $0.00 |
| Cache write, 5 min | 210.3K | $6.25 | $1.31 |
| Cached input | 986.9K | $0.5 | $0.49 |
| Output | 31.1K | $25 | $0.78 |
Opus 5
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 262 | $5 | $0.00 |
| Cache write, 1 hour | 328.4K | $10 | $3.28 |
| Cached input | 22.6M | $0.5 | $11.28 |
| Output | 66.3K | $25 | $1.66 |
Run total: $18.81 API-equivalent, $1.07 estimated subscription. 24.2M tokens; 4.7% of a week’s allowance.
GLM 5.3 via Claude Code
GLM 5.3
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 279.8K | $1.4 | $0.39 |
| Cached input | 18.2M | $0.26 | $4.73 |
| Output | 118.7K | $4.4 | $0.52 |
Run total: $5.64 API-equivalent, $0.24 estimated subscription. 18.6M tokens; 1.9% of a week’s allowance.
GLM 5.3 via ZCode
GLM 5.3
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 143.2K | $1.4 | $0.20 |
| Cached input | 15.2M | $0.26 | $3.95 |
| Output | 53.3K | $4.4 | $0.23 |
Session average: $4.39 API-equivalent, $0.19 estimated subscription. 15.4M tokens; 1.5% of a week’s allowance.
Codex GPT-6 Astra low
GPT-6 Astra
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 105.8K | $10 | $1.06 |
| Cached input | 4.1M | $1 | $4.08 |
| Output | 18.3K | $50 | $0.92 |
Run total: $6.05 API-equivalent, $0.17 estimated subscription. 4.2M tokens; 0.4% of a week’s allowance.
Sonnet 5 medium
Opus 4.7 automatic review
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 35 | $5 | $0.00 |
| Cache write, 5 min | 121.2K | $6.25 | $0.76 |
| Cached input | 719.6K | $0.5 | $0.36 |
| Output | 16.3K | $25 | $0.41 |
Sonnet 5
| Token type | Tokens | $ / 1M | API $ |
|---|---|---|---|
| Fresh input | 364 | $2 | $0.00 |
| Cache write, 1 hour | 471.5K | $4 | $1.89 |
| Cached input | 44.4M | $0.2 | $8.89 |
| Output | 126.6K | $10 | $1.27 |
Run total: $13.57 API-equivalent, $0.78 estimated subscription. 45.9M tokens; 3.4% of a week’s allowance.
The subscription assumptions
| Plan | Monthly fee | Monthly API-equivalent allowance |
|---|---|---|
| Claude Max · Fable | $100.00 | $1,875 |
| Claude Max · Opus / Sonnet | $100.00 | $1,750 |
| GLM Pro annual | $56.00 | $1,300 |
| Codex Pro | $200.00 | $7,000 |
These are the subscription comparison’s existing allowance assumptions. The Claude $100 estimates are one quarter of the measured $200 plan allowances. GLM’s $1,300 estimate came from three weekly exhaustion windows. Codex uses the comparison’s central case.
The usage audit did not remeasure those allowances. These costs allocate a subscription fee; they are not per-task charges. Weekly share is the estimated subscription cost divided by the monthly fee, multiplied by 4.33.
One feature, six configurations
The task was to rebuild Bruno’s public chat bot so a price question runs the real solar estimator in conversation form. The configurations started from the same base commit in isolated worktrees. For ZCode, cost and token usage are the arithmetic mean of two sessions: $0.21 and $0.17 become $0.19 per session. The 8.75 grade and recorded wall time describe their shared final branch, not independently graded attempts.
Fable and Opus ran at low effort, Sonnet at medium, and Codex Astra at low. GLM used Claude Code’s default thinking and ZCode’s max reasoning variant. These are the configurations tested, not identical reasoning budgets.
One Fable 5.1 reviewer graded the branches with file-level evidence and re-ran the chat tests. Each rubric item scores 0 to 2; the total is divided by 2 for a maximum of 10. Self-reports were treated as claims, not evidence.
The 10-item rubric
- Intent gate catches all 7 openers in both languages
- Step machine, parsers, skips known data
- Reuses the wizard's geocoder, numbered pick of at most 3
- 'no sé' falls back to the published sample size
- Lead and report through the real libraries, correct money
- Prompt cleanup: no price dump, no early visit, 'nosotros'
- Interruption answered, then the pending question repeated
- Tests: openers, happy path, unknown consumption, interruption, ambiguity
- Gate green, lane respected, trailer, not merged
- Real end-to-end transcript and an honest report
The code under test is private, so the exact task cannot be re-run from the public repository. The redacted prompt, grading evidence, reports and extraction scripts are public. One feature evaluation does not establish a universal model ranking.
The cost correction
The first version double-counted repeated assistant-message records, priced one-hour cache writes as five-minute writes, and priced automatic review calls as the main model. This version counts each message once, uses the correct cache duration and prices each model separately.
GLM via Claude Code uses its final aggregate usage summary. ZCode averages the unique provider request totals from its two sessions. Codex uses its final cumulative turn usage, with cached input separated from fresh input. Anthropic’s list prices supply the model and cache rates.