Result: no convincing speed benefit from low effort in this sample. Both efforts passed every task and check. Low reduced aggregate elapsed time by just 0.39%, and was faster in only 6/15 matched task/repeat pairs. Its median paired difference was 1.026s slower. One medium inventory run (123.023s versus 82.262s at low) accounts for much of the aggregate difference. Do not change the default on this evidence alone.
Setup
- Model:
openai:gpt-6-astra, using configured ChatGPT subscription credentials / Responses API. - Five repeats per task and effort; three tasks; 30 total attempts. Each starts in a fresh workspace, with the same prompt and independent grader.
- For each task, order was medium → low in odd repeats, low → medium in even repeats. Tasks stayed in the same order. Runs were sequential; cache warmth and provider load were not controlled.
- Effort was explicitly passed to the provider in memory. Saved settings were not changed. Other megacode settings/instructions remained configured defaults or user settings; no new isolation of global instructions was introduced.
- Automatic approvals, 240s internal abort, 300s process deadline. No process failures, aborted requests, or timeouts occurred. Tool-level nonzero shell results occurred during log recovery; these are retained in the raw metrics and did not prevent passing the grader.
- Baseline rerun: 0/3 tasks and 1/15 checks, as expected.
- This is a fresh run of the current instrumented code, not a direct reproduction of the older single-run harness comparison. Claude Code/Codex were not rerun in this experiment.
Aggregate results
| Metric | Medium | Low |
|---|---|---|
| Tasks passed | 15/15 | 15/15 |
| Checks passed | 75/75 | 75/75 |
| Total process wall time | 1,076.706s | 1,072.491s |
| Mean task wall time | 71.780s | 71.499s |
| Median three-task repeat wall time | 214.666s | 214.447s |
| Total model-request time | 1,066.961s | 1,062.970s |
| Model-request share of process time | 99.09% | 99.11% |
| Total tool time | 2.628s | 2.598s |
| Unattributed wall time (startup, logging, etc.) | 7.117s | 6.923s |
| Model steps | 105 | 108 |
| Tool calls | 93 | 100 |
| Reported input tokens | 236,406 | 232,739 |
| Reported output tokens | 20,466 | 20,192 |
| Median first-visible-text latency (text-emitting steps only) | 3.291s | 3.211s |
| Steps with visible text / all steps | 29/105 | 25/108 |
Model-request time includes preparation, streaming, provider retries and text callbacks; it is not pure server inference time. First-visible-text latency excludes reasoning/tool-argument streams and is null for tool-only responses, so it must not be averaged as zero. Token accounting is the adapter’s existing input/output accounting; cache usage was not separately recorded.
Per-task latency distribution
Five attempts per cell. Times are full process wall seconds; all attempts succeeded.
| Task | Medium mean / median / min–max | Low mean / median / min–max |
|---|---|---|
| range-parser | 56.349 / 58.094 / 44.763–66.044 | 61.068 / 62.980 / 55.464–67.141 |
| inventory-transaction | 93.598 / 88.567 / 81.623–123.023 | 88.982 / 88.652 / 82.262–96.914 |
| log-recovery | 65.395 / 66.955 / 61.137–69.513 | 64.448 / 63.397 / 57.290–75.093 |
Every attempt
Execution order is retained. Durations are seconds; per-model-step and per-tool-call measurements are in the JSON snapshot.
| Repeat | Task | Effort | Checks | Wall | Model | Tools | Steps | Calls |
|---|---|---|---|---|---|---|---|---|
| 1 | range-parser | medium | 5/5 | 58.094 | 57.466 | 0.141 | 7 | 6 |
| 1 | range-parser | low | 5/5 | 55.889 | 55.359 | 0.126 | 7 | 6 |
| 1 | inventory-transaction | medium | 6/6 | 91.419 | 90.902 | 0.135 | 7 | 6 |
| 1 | inventory-transaction | low | 6/6 | 96.914 | 96.215 | 0.139 | 9 | 8 |
| 1 | log-recovery | medium | 4/4 | 69.513 | 68.760 | 0.228 | 7 | 9 |
| 1 | log-recovery | low | 4/4 | 57.290 | 56.490 | 0.241 | 7 | 6 |
| 2 | range-parser | low | 5/5 | 63.867 | 63.195 | 0.137 | 6 | 6 |
| 2 | range-parser | medium | 5/5 | 62.995 | 62.408 | 0.145 | 7 | 6 |
| 2 | inventory-transaction | low | 6/6 | 88.652 | 88.114 | 0.145 | 7 | 6 |
| 2 | inventory-transaction | medium | 6/6 | 83.356 | 82.634 | 0.122 | 7 | 6 |
| 2 | log-recovery | low | 4/4 | 63.397 | 62.503 | 0.278 | 8 | 7 |
| 2 | log-recovery | medium | 4/4 | 62.371 | 61.539 | 0.267 | 8 | 7 |
| 3 | range-parser | medium | 5/5 | 44.763 | 44.066 | 0.157 | 5 | 4 |
| 3 | range-parser | low | 5/5 | 55.464 | 54.766 | 0.140 | 7 | 6 |
| 3 | inventory-transaction | medium | 6/6 | 88.567 | 88.037 | 0.133 | 7 | 6 |
| 3 | inventory-transaction | low | 6/6 | 86.326 | 85.797 | 0.142 | 7 | 6 |
| 3 | log-recovery | medium | 4/4 | 61.137 | 60.311 | 0.260 | 8 | 7 |
| 3 | log-recovery | low | 4/4 | 65.749 | 65.128 | 0.227 | 7 | 9 |
| 4 | range-parser | low | 5/5 | 67.141 | 66.635 | 0.108 | 7 | 6 |
| 4 | range-parser | medium | 5/5 | 49.847 | 49.165 | 0.140 | 7 | 6 |
| 4 | inventory-transaction | low | 6/6 | 82.262 | 81.734 | 0.138 | 7 | 6 |
| 4 | inventory-transaction | medium | 6/6 | 123.023 | 122.499 | 0.134 | 7 | 6 |
| 4 | log-recovery | low | 4/4 | 75.093 | 74.442 | 0.260 | 8 | 7 |
| 4 | log-recovery | medium | 4/4 | 66.955 | 66.267 | 0.219 | 8 | 7 |
| 5 | range-parser | medium | 5/5 | 66.044 | 65.477 | 0.141 | 7 | 6 |
| 5 | range-parser | low | 5/5 | 62.980 | 62.458 | 0.130 | 7 | 6 |
| 5 | inventory-transaction | medium | 6/6 | 81.623 | 81.081 | 0.153 | 5 | 4 |
| 5 | inventory-transaction | low | 6/6 | 90.755 | 90.228 | 0.134 | 7 | 6 |
| 5 | log-recovery | medium | 4/4 | 66.999 | 66.348 | 0.255 | 8 | 7 |
| 5 | log-recovery | low | 4/4 | 60.712 | 59.908 | 0.251 | 7 | 9 |
Interpretation and next optimization
Local execution is not the bottleneck here. Even eliminating all measured tool time would save only about 0.24%. Eliminating all non-model overhead would save under 1%. Parallelizing local reads alone cannot close the earlier Claude Code gap.
Prioritize fewer model round trips, then measure latency on the same model/backend. Low effort produced slightly more model steps (108 vs 105) and tool calls (100 vs 93), which may offset any per-request benefit; this sample does not establish causation. Keep medium as the default pending broader correctness and latency measurements.
Reproduce and inspect
npm run bench -- --model openai:gpt-6-astra --label megacode-effort \
--effort medium --effort low --repeats 5
- Checked-in metrics snapshot: all 30 results, including each model/tool timing, no transcripts or credentials.
- Raw artifacts (gitignored):
.bench/2026-10-03T14-30-34.388Z-megacode-effort/, including workspaces, transcripts, prompts, grader output, and metrics. - Baseline:
.bench/2026-10-03T14-58-57.086Z-baseline/results.json.