The methodology

How software gets built when agents do the work.

Software companies have to adapt. The six phases we all learned still hold, but every one of them assumed a person did the work. This is how LUM keeps the agents productive and the code trustworthy at the same time: one framework, roles with hard limits, and a human at every gate.

01 · The tension

The hard part is the balance.

Code got cheap. An agent can write more of it, faster, than any team you could hire, overnight, for a fraction of the cost. That single fact splits every software company in two.

Refuse the agents, and you cannot compete on price or speed. Unleash them, and you drown in code nobody can maintain. Neither side wins on its own. The firms that last will be the ones that find the balance, productive agents and trustworthy code, and make it repeatable.

When code is cheap, the problem moves to three things.

02 · The approach

Three questions a model can't answer.

Judgment. What is worth building, and is this the right thing? An agent will build whatever you point it at. Deciding what deserves to exist, and why, is human judgment. It opens every loop, in the Analysis phase.

Control. What may an agent touch, change, and spend? Set once per client, when they come on board: which agents may work, what each may change, how much time and money it may spend, and who approves a release. Rules the agent cannot step over, enforced by tools, not prompts.

Quality. How does code stay readable, fixable, and trusted? One stack, a plan and tests before any code, and checks that fail the work the moment it leaves the lines. So a person can pick up any part and fix it, with or without an agent.

03 · The loop

Six phases, closed into a loop.

Analysis, design, development, testing, deployment, maintenance, the loop every engineer knows. In LUM it runs continuously. A person shapes an intent and writes the spec; then the plan, build, test, and review run between sessions; a person approves the plan, approves the pull request, and promotes to production; the observer watches, and findings feed the next loop.

Everything mechanical happens unattended and leaves evidence. The team's hours go to judgment, the three things above, and nothing else.

04 · The roles

Each agent does one job, with its own scope.

Five roles. Each has a model class, a set of tools, and hard caps on time and spend. None can approve its own work, merge to main, or widen its own budget.

RoleModelWhat it does
PlannerbestGiven a spec, it writes the plan, the interfaces, and the tests that name the behaviour, before any code. A second reviewer critiques the plan before it is committed.
BuilderworkerGiven a clear plan, a bounded module, and tests that name the behaviour, it delivers working code overnight, filling in the bodies until the tests pass.
ReviewerbestReviews the pull request against the principles, from a different vendor than the one that wrote the code. Its check is required alongside a human's.
ObserverutilityWatches production all night. Given logs and an alert, it finds the cause, drafts the fix, and hands the judgment call back to a person.
ExplorerutilityRead-only research for another role. Returns an answer with references; writes nothing.
05 · Authority

Rules the agent cannot step over, enforced by tools, not prompts.

Authority does not live in a prompt an agent could talk its way around. It lives in the tools: branch protection, a hook that denies edits to protected paths, a check that fails a diff outside the plan's module, one secret per run, hard caps.

Its own branch only

Never main, never a force push. A human approves every merge.

No self-authorisation

It cannot approve or merge its own work, or change the rules, the review file, or its own budget.

One secret per run

No production data, no deploys, no client contact.

Caps end the run

Time, turns, and spend are bounded; when a cap is hit, the work is pushed with a note and a person is told.

Autonomy widens by evidence

One rule at a time, earned, never assumed.

06 · The framework

One stack. Tests before any code. Checks that fail fast.

Sixty-plus years of combined time building production software, poured into guardrails the agents cannot step over. Every project has the same shape, one bounded, vetted stack; a shared library of capabilities behind interfaces; a project template; written principles, so people and agents move between projects without relearning.

The result is code a person can read, fix, and trust. Explicit over clever. No magic. Any team member can hotfix anything, with or without an agent. That is the bar the checks defend on every pull request.

07 · Learning

Every finding becomes a rule at the cheapest level that holds it.

A mistake is only worth making once. When a review or an incident turns up a lesson, it enters at the lowest rung that can hold it, and never higher. The system improves; the instructions stay short.

Rung 1

A lint rule or hook

When it is mechanical, a tool enforces it. No one has to remember.

Rung 2

A test

When it is a behaviour to keep, a failing test guards it forever.

Rung 3

A line in one role

When it concerns one job, it goes in that role's or specialisation's instructions.

Rung 4

A shared principle

Only when it applies to every run, and only after it has happened twice.

08 · The platform

One place from which to run the company.

Every client, the projects each one has, the health of each project, and what every agent is doing, on one screen. Every intent is a card that moves through the loop on its own as agents plan, build, test, and release. A human approves from the board and reads the state of a whole project at a glance.

Approvals, budgets, and the chain from intent to release live in the platform. What an agent may touch is set once per client, when they come on board. And when there is nothing to decide, the platform says so: nothing else needs you, while nine agents work across seven projects.

09 · A worked example

A fifteen-minute intent. A pull request while you sleep.

A clinic asks for a report of patients whose import email differs from the portal email: flag them, do not merge them. The tension is in the detail, and the tests hold it.

session · one intent, end to end
you  Also flag any patient with two different emails.
planner  Good. Read model and existing dedupe. Listing those for the reviewer.
system  Six new tests fail as expected.
builder  Implementing the matcher now.
system  Dry run on the 41k-row sample. Not merging them.
observer  The cause and a drafted fix are ready. The judgment call is yours.
10 · The result

A software platform built from the ground up on this methodology.

Not a tool bolted onto an old process, the process itself, rebuilt for a world where code is cheap. Productive agents and trustworthy code, at a price a business can actually pay, with the code, the context, and the reasoning owned by you.

11 · The studio

The IP, put to work for clients.

This methodology is not only how the platform runs; it is what LUM Studio sells. The Studio is the implementation arm: consulting and custom software development where every engagement inherits the framework, the agent roles, and the human gates, from a net-new build to taking over a legacy system to installing the guardrails in your own team.

The two arms feed each other. Studio engagements fund and sharpen the platform; the platform makes every Studio engagement faster and more trustworthy than a consultancy starting from zero.