Quaedra Research

megacode

A minimal coding agent for the terminal. One harness for Anthropic, OpenAI, Gemini and any OpenAI-compatible model, local ones included.

v0.7.0Latest release, MIT
9Providers, live model lists
75/75Benchmark checks passed

How it works

megacode runs a plain agent loop: a model turn, then its tool calls, repeated until the model stops. It talks to each provider through its official SDK and keeps the conversation in a neutral format, so you can switch models mid-conversation. Each turn also keeps the provider's own content, so thinking blocks and thought signatures go back to the same provider unchanged.

Tool results are kept small. File reads return 200 lines at a time, long command output keeps its first and last parts with the full log saved to disk, and the conversation is summarized when it nears the context window.

FeatureWhat it does
ProvidersAnthropic, OpenAI (or a ChatGPT plan), Gemini, OpenRouter, Groq, DeepSeek, Ollama, LM Studio, any compatible server
ToolsRead, write and edit files, bash, grep, list files, view images, ask questions
MCP serversstdio, HTTP and SSE, in the same config format as Claude Code
SkillsInstall SKILL.md skills from a folder or GitHub
WorktreesWork in a separate git worktree with -w
SessionsEvery conversation is saved; continue it with --resume

Written in TypeScript with an Ink terminal UI. The core has no SDK, file or UI dependencies; providers and tools plug in as adapters. Layout

Benchmarks

The harness benchmark is a small smoke test: three fixed coding tasks (a parser, an atomic inventory update, and finding a bug in long log output), each in a fresh workspace and graded by 15 checks the agent never sees.

Same model, two harnesses

Mean wall time per task in seconds, lower is better. Both on GPT-6 Astra at medium effort through the same gateway, five repeats per task, alternating order.

HarnessChecksSuite timeInput tokensTool calls
megacode75/75219.3 s264,81898
Claude Code75/75206.2 s172,11045

Suite time is the mean of the five three-task runs. Input tokens include cached context and are not cost. Fifteen attempts each; Claude Code 2.1.288 in bare mode. October 3, 2026. Every attempt

Both passed every check. Claude Code was 5.9% faster in total and faster in 11 of 15 paired attempts, using about half as many tool calls. 99% of megacode's time is spent waiting on the model, so fewer round trips is where the gap is.

Input tokens vs Codex

Input tokens per task, including cached context, lower is better. Both on GPT-6 Astra at medium effort via a ChatGPT plan, one run each.

Single runs on October 2, 2026; cache warmth wasn't controlled. Both passed 15/15 checks; megacode took 188.4 s in total and Codex 230.3 s. Details

With the same model, megacode sent 80% fewer input tokens than Codex. That's one run per harness, not a general claim about speed or cost.

Reasoning effort

Low effort didn't make megacode faster: across 30 attempts it saved 0.39% of total time and was faster in only 6 of 15 pairs, with every check passing at both levels. Results

Run it

Needs Node 22.14 or later. Run /login to connect a provider; ChatGPT and OpenRouter support browser sign-in.

npm install -g @megacode/cli

megacode                          # interactive
megacode "fix the failing test"   # one-shot
megacode -m ollama:qwen3:8b       # any provider:model
megacode -w fix-auth              # in a new git worktree

Read the docs for providers, settings, skills, MCP servers and sessions.