Home > Blog > Fast is easy. Provable is the product.

Fast is easy. Provable is the product.

August 19, 2026 — 19 min read

Why we run AI delivery in Human-Agent Pods: headless agents, deterministic quality gates, and a named human signing every artefact.

Somewhere in your delivery organisation, right now, an engineer is pasting part of your codebase into a model you did not choose, under terms you have not read. Perhaps you have governed this. Enterprise Copilot seats, an allow-listed model, a network proxy, DLP rules, usage telemetry. We concede all of it; it is real work and it matters.

It is also governance of the wrong object. Those controls govern access to the tool. They cannot tell you which model touched which artefact, what data crossed the boundary on a given task, what quality bar the output cleared, or who signed the result. They govern the tool. They do not govern the work.

That is why we hold that a copilot inside an engineer’s IDE is ungovernable by construction: governance that depends on a developer choosing to comply is not governance. Our answer was to move the AI off the engineer’s desk, into a governed harness that runs only approved tools and models, and to move the humans up to the gate, where their job is to decide and answer for the work rather than to type it. We call the resulting unit a Human-Agent Pod. This piece explains what one is and shows you a production agent in enough detail to judge.

A pod is people who sign, paired with agents that produce


A Human-Agent Pod is a delivery team in which every AI agent is paired with a named, accountable human, running a spec-first SDLC under a control plane we call AI PodOps. The agents produce artefacts. The humans apply expertise to judgment: they steer, they decide, and they sign. Nothing proceeds without that signature.

We run pods in two shapes.

A mirrored pod pairs one agent with one specialist: BA, solution architect, PM, fullstack developer, UX/UI designer and QA, each with their own agent. That buys dual control, different signatories at every stage, parallel execution and a complete audit trail, which is why we point mirrored pods at regulated sectors: banking, insurance, healthcare, the public sector.

A compressed pod puts one senior generalist over two or three agents: a tech lead over the SA and backend developer agents, or a senior BA/QA lead over the BA, PM and QA agents. Lower day rate, faster handoffs, one person holding the whole picture. That shape fits MVPs and internal tools.

Across both shapes, our commercial model targets up to 40% lower blended cost than the equivalent all-human team at the same scope. We say “targets” deliberately: that is a claim about the pod model, not a measured average, and the audit trail is what lets a client hold us to it.

One rule is absolute: one agent, one accountable human. If you run delivery on RACI, read the pod as a deliberate rewrite of one letter. Responsible moves to the agent: it produces the artefact, escalates when it is blocked, and reaches out when it needs a decision.

Accountable stays with the human: they judge, they sign, and they answer for the work by name. Accountable never moves. In our staffing view, an agent without an assigned human is not an efficiency gain. It renders as uncovered risk, in the same red as an unfilled seat.

Where this sits in a crowded market

If you are evaluating AI delivery in 2026, you are being offered one of three things, and all three are legitimate. They are just not this.

The platform suites (GitHub’s Agent HQ, ServiceNow’s AI Control Tower, Atlassian’s Rovo) govern agents against their own operating surface, and that governance is real; two of the three now reach third-party agents, and we will not pretend otherwise. Their anchor is still the platform: the repository, the CMDB, the suite. What none of them holds is the delivery workflow itself: the chain from spec to release, a deterministic quality bar at each stage, and a named human signing each artefact, wherever the work happens to run.

The autonomous coders (Devin, the Cursor agent family) sell capability, by now with serious enterprise trimmings: SOC 2 attestations, access controls, audit APIs. But what their logs record is who ran an agent and when, not which approved spec a change traces to, which deterministic checks it cleared, or which named human answered for it at a gate. The operating model is delegation to autonomy, run by an engineer, and a non-engineer can neither operate the work nor reconstruct it as evidence.

The frontier labs sell enterprise capability deals: excellent tooling for a company that wants its own engineering org to move faster, where the client still runs it, governs it, staffs it and answers for it.

A pod is the fourth thing: a cross-tool agent-human team under one harness, operable by a non-engineer. We govern the pod across the tools it touches and keep the delivery evidence: a system of record for AI delivery governance.

Which also means the frontier labs are our suppliers, not our competitors. The platform is model-agnostic by design: agents are swappable by contract, every agent’s model plane is disclosed, and local models can run inside the client’s tenant. The labs sell the engine; we show up with the car, the driver, and the insurance, and we buy engines from whoever makes the best one. The honest boundary runs the other way too: if what you want is your own engineering org moving faster, buy the lab deal. A pod is for buying delivered outcomes with proof attached.

The chain is the dependency graph

Every pod runs the same artefact chain: spec, design, UX/UI spec, tasks, code, review, test, release, with a Knowledge Base agent running pod-wide beside the chain rather than inside it. Each stage consumes the artefact before it and produces the artefact after it, and that produces-to-consumes relation is not documentation of a process. It is the process: the dependency graph the platform executes.

Sheet showing role, consumes-produces and artefact list.

Two properties fall out of taking the graph seriously.

Entry is always the chain’s first node. Every confirmed piece of work lands at the BA, so every downstream artefact traces back to a spec. When a client arrives with a spec of their own, the BA agent adopts it: ingests it, runs the validators against it, conforms it to the template or flags the gaps. The author becomes the inspector, and traceability starts in the same place either way.

Rework travels the chain. When QA fails a build because the design was wrong, the rejection targets the design stage, the root cause, and not merely the developer one step back. Because produces-to-consumes is the dependency graph, everything downstream of a corrected artefact is stale by definition and re-runs forward; no human has to remember what a corrected design invalidated.

Compare a human team, where a corrected design document quietly coexists with code built on the old one until an integration test, or a customer, finds the seam. The economics view attributes the rework cost to the stage that caused it, and repeated bounces escalate to the accountable human in Slack.

One project, one memory, and an agent that owns it


The Knowledge Base agent is the pod’s cornerstone, and it is mandatory in every pod rather than a toggle a client can switch off. It holds the project’s truth: legal and contractual constraints, decisions with their rationale, open conflicts, the glossary and domain sources, task context. Every agent and every human draws on the same context, visibly and auditably. That shared memory is the difference between a pod and six chatbots, each hallucinating its own version of the project.

It is a curator, not a bucket. The memory is organised as a knowledge graph, with Graphiti as the cornerstone of how the data is structured, and it connects to wherever a project’s truth actually lives: email, Google Drive, Notion, Slack, Teams, HubSpot, Teamwork, Jira, and any other source a project uses to collect and manage information.

As it ingests, it curates. And when two sources contradict each other on something it cannot resolve itself, it does what every agent in the pod does with a decision above its pay grade: it escalates to a human, who decides which claim becomes the project’s truth. Even the memory has a gate.

The consequence a delivery lead feels first: the statement of work lives in that shared context as a source of truth, and produced work is checked against it, with deviations flagged. “Is this a bug or a change request?” stops being a negotiation between a PM and a client and becomes a lookup with a citation. On fixed-price work, that single property can be worth more than the delivery speed.

The honest limit: the pod knows only what it has been given. Ingestion scope, redaction, and the client’s right to ask “what does it know, and can I remove this?” are operational questions we design for, not footnotes.

Anatomy of a real agent: our BA

Diagrams are cheap. So here is one of our agents in production, in enough detail to judge whether this is engineering or a slide.

The BA agent is Python 3.12, FastAPI and LangGraph. A ReAct supervisor is the single entry point; the spec-writing pipeline is a tool the supervisor calls, never invoked directly. Work arrives when a ticket moves into the agreed column in the client’s tracker, Jira or Teamwork, which fires a webhook to the agent’s trigger endpoint, fire-and-forget, and the supervisor loop starts in the background.

From there, the supervisor can search the tracker task with its comments and attachments, search a pgvector project memory of domain knowledge and past spec patterns, write the spec, run a consistency check of the draft against every previously approved spec (data models, APIs, business rules, terminology), share the draft in Slack, wait for a reply, and record approval, with its Slack tools loaded dynamically over MCP.

The pipeline itself is a bounded loop: generate, evaluate, route, then at most two refinement turns. It runs under a ten-minute timeout and a five-dollar hard cost cap per run, and the best spec across turns is tracked and restored if the final turn regresses.

Now the part we most want you to understand: the evaluate node makes zero LLM calls. The spec is graded by eight structural validators, token-level checks over ids, counts, set equality and presence. One is representative: it checks that the set of AC-NNN acceptance-criterion ids in the plain-language section exactly equals the set in the EARS technical section, which catches orphaned criteria on either side.

Another enforces EARS syntax on the acceptance criteria themselves (“When”, “While”, “Where”, “The system shall”) and fails the run below 80% coverage.

A third checks that every low-confidence decision in the decision log carries a substantive rationale, at least twenty characters and not a TBD. The other five are the same species: presence, parity and duplicate checks over sections, ids and message codes. On top sit six completeness scores (user roles, business rules, acceptance criteria, scope boundaries, error handling, data model), starting at 95 and capped down by failed validators; a run passes only when all six hold at 80 or above.

Why insist on a deterministic grader? Because a model evaluating a model tells you what a model thinks, and the design premise of the whole system is that the AI does not grade its own homework. But be clear about what the validators prove: that a spec is well-formed, not that it is correct. A spec can pass all eight checks and still describe the wrong product. Structural quality is machine-checkable; semantic truth is not. That is exactly why a human still signs it.

The signature is engineered too. When the agent needs a clarification, it calls LangGraph’s interrupt() and the graph stops; a Slack Events webhook resumes it with the stakeholder’s reply. Approval is bound to an authorised Slack user id configured per project, not to whoever answers fastest. Every spec version is stored with its scores and its source (generated, refined or manually edited), diffs between versions are viewable, and the decision log surfaces a low-confidence count so the reviewing BA knows how many assumptions to probe.

By our own run telemetry, a production spec run costs roughly $0.08 to $0.17 in model spend. Put that next to a human afternoon. And the unglamorous edge is engineered as well: post-approval chores (move the tracker column, post the comment, upload the spec file) run with retries and a dead-letter queue, because integrations, not reasoning, are where agent systems actually die.

Taking implementation headless

The BA proves the pattern. Implementation is where the thesis bites, because implementation is where copilots live today.

Our Developer and QA agents run the same way: Claude Code in headless mode, on our own server. No IDE, no engineer’s laptop, no personal account. A harness we control, running only the tools and models we have approved, on infrastructure where data residency is a configuration rather than a hope.

Today the runner draws on a Claude subscription rather than per-token API metering, which changes the economics of iteration: the marginal cost of another implementation attempt is predictable instead of a variable that punishes rework. We hold that arrangement honestly. Headless mode is an officially supported way to run Claude Code; every seat belongs to one accountable human; and Anthropic has signalled that heavily shared production automation belongs on API billing, so we treat the current pricing as a commercial term that can change, not a structural advantage. If the meter changes, the meter changes. The harness and the governance around it do not.

The harness is not married to a single engine, either. Claude Code can be pointed at local models running on our own H200 cluster, large open-weight models such as Qwen or DeepSeek, where a client’s data must never leave the tenant or where the economics favour it. That works better than it might sound, and the chain explains why: by the time implementation starts, the hard thinking has already been done and signed upstream, in the spec, the architecture and the task breakdown. The runner receives a small, micromanaged context per task, so the model executing it does not need to be the strongest one on the market. Most of the intelligence in the system lives in the artefacts, not in the model of the week.

The runner does not read the ticket. It reads the chain: the approved SPEC.md, ARCHITECTURE.md, UIX-UI-SPEC.md and TASKS.md, the project CONSTITUTION.md (stack, coding, testing, API and security standards, and the quality-gate commands), and the Knowledge Base agent’s shared context. That is the complete brief, and it exists because every upstream stage was gated.

For visual work, web or mobile, one more gate runs before a single line of code. Our own Figma MCP takes the wireframe, produces a rendering of the feature without writing the code, compares the rendering against the Figma design, and escalates if the match is not 100%. Development does not start until it is, because we know where the time actually goes: fine-tuning the last 5% of a UI after the code exists is how implementation becomes the bottleneck. We settle the pixels before the first commit.

Execution runs in three phases. Plan: decompose the work against TASKS.md and isolate it in a git worktree. Implement: test-driven development as an iron law, per task, red then green then refactor, and code written before its test is deleted. Before anything is offered for review, a blocking checklist must pass: suite green, lint clean, coverage met, scope integrity intact. Review: a code-review pass that is forbidden from approving without reading the code and running the tests, with bounded auto-fix attempts before it must escalate to a human.

What comes back is not a chat transcript. It is artefacts: an IMPLEMENTATION-SUMMARY.md, a REVIEW.md, a branch and a pull request. Those artefacts are what the human gate reviews and what the ledger records.

The QA agent works the same way, one layer up. Under the hood it drives Playwright; over it sits our harness. Test automation usually dies on its inputs: vague acceptance criteria and undocumented flows. Here the inputs are the chain itself: EARS acceptance criteria, the UX/UI spec, the implementation summary, each validated and signed before QA ever runs. We are automating the automation, with artefacts prepared for exactly that.

One deliberate inversion is worth explaining. The spec gate lives in the platform, because a PM should never be dropped into engineering tooling to do their core job. The code gate is the GitHub PR review itself, because engineers review code in GitHub and moving them anywhere else would be theatre. The webhook resolves the gate; the platform records it. Meet each persona where they already work.

Provenance lives in the repository, and not only on a dashboard. The Dev agent commits under a bot identity with signed commits and a Co-authored-by: line naming the accountable human, and every PR carries a contract in its description: the spec link, the gate id, the validator summary. Six months from now, git blame answers the question your auditor will ask.

Where the human actually sits

Lift the human out of the IDE and their job changes shape. It stops being production and becomes judgment: reading what the agent produced, applying their expertise, deciding, and answering for the decision. The pod is built so that the job happens where people already work.

Agents reach humans in Slack, proactively. Clarifications, approval gates and escalations arrive as messages, and people answer them the way they answer a colleague. The dashboard is optional depth, one deep link from any message to the exact decision, never a required habit, because a gate that only exists on a dashboard nobody opens is not a gate. Approval authority is bound to configured user ids rather than to whoever replies first, and every message in either direction lands in the audit trail.

Here is what a gate actually looks like. At 09:12 the BA agent hits an ambiguity in a payment flow, calls interrupt(), and stops. A Slack message reaches the accountable BA: the draft spec, the low-confidence count, the specific question. They read for four minutes and reject with a typed reason: “refund path contradicts BR-014; confirm with the client before AC-031 stands.” The graph resumes with their answer. The ledger records the spec version, the validator results, the reason and their user id. That is the entire ceremony, and it is also the entire audit story.

Typed reasons are required on every rejection and every override. Clean approvals offer structured quick-reason chips instead, so the ledger never fills with forty rows of “Looks good.”

Autonomy is earned, never assumed. The ladder runs from L0 (review everything) through batch review and one-in-five sampling to auto-clearing low-risk items, and promotion criteria are deterministic validator streaks. The system proposes, with evidence of the form “BA eligible for L2: 47 consecutive 8/8 specs, 4% rejection rate”; the accountable human grants; the ledger records. Downgrade is free; upgrade is gated. Autonomy you can audit your way into.

And because agents run around the clock while people do not, human capacity is modelled as a first-class object: working hours, out-of-office status, named delegates, gates-per-day capacity, a coverage timeline where an uncovered hour flags the same red as an unstaffed agent, and an intake throttle that pauses upstream work when someone’s open-gate queue exceeds their capacity. Coverage is accepted with a one-tap audited acknowledgement. It is never silently assigned.

Proven where it counts

None of this is a concept deck. The BA agent runs in production federation, with real runs, real gates and real health reporting behind the numbers in this piece. The Knowledge Base agent and the append-only audit ledger operate alongside it. The Spec-First workflow that disciplines the whole chain, six gated steps and eleven mandatory artefact templates, is in daily use on client and internal projects across Cursor, Claude Code, OpenCode and Codex. The pattern you have just read is the pattern we run.

Nor is it a pivot. Q has delivered full-team SDLC engagements for years: whole teams, spec to release, under one roof. The pod is a multiplier on that muscle, not a replacement for it, and the chance to run delivery the way we always wanted to: gated, traceable, provable at every step.

One limit we will state ourselves, because it is the reason the model has humans in it at all: structural validation cannot catch a spec that is well-formed and wrong. We narrow that gap with a semantic layer on top of the deterministic one: LLM-driven consistency checks that compare each new spec against everything already approved and flag duplicated work, overlaps and contradictions as warnings and guidance while the spec is being generated. But warnings advise; they do not decide. Only the person at the gate signs. That is not a gap in the system. It is the job description.

“Show me”

Every serious conversation about AI delivery eventually arrives at the same request: not “how fast,” but “show me.” A copilot cannot answer it, because the evidence never existed. A pod answers it by construction, because the evidence is a by-product of how the work moves. Not an autonomous coder you must trust: a governed agent-human pipeline you can prove.

Part 2 takes the harder half of the question: what it means to launch, run, and recover a pod on a real engagement, around the clock, with a non-engineer on coverage. The central dashboard, the AgentOps underneath, incidents and recovery, the economics, and the regulatory evidence a client can hand to their auditor.

Zlatko Matokanovic
Zlatko Matokanovic

Q's Director of AI Research and Development has over 11 years of experience in mobile development, mainly in iOS. Sharpening his organizational and people skills while leading a new vertical at Q are his forte, and when not working, Zlatko loves to play drums, tinker with IoT, do crossfit, and try to make every day interesting for his two young sons.

GIVE KUDOS BY SHARING THE POST!

Partner with us