novianworks Get in touch

Blog

Progressive Determinization

What if the first version of an app was just a prompt, and you swapped it for real code a piece at a time?

If LLMs were always right, instant and free, a lot of software would collapse into a single step: say what goes in, say what should come out, and let the model handle everything in between.

They aren’t. They make things up, give different answers to the same question, take seconds to respond and cost money on every call. Most of the discussion about building with LLMs is about working around that.

I’ve been looking at it the other way round. Those weaknesses are useful information: they point at the parts of a program that should be ordinary code. I call the approach progressive determinization. The name is borrowed from automata theory, where determinization turns a nondeterministic automaton into a deterministic one. The analogy is loose, but the direction is the same.

Start with a prompt

Normally you’d begin with an architecture: a domain model, services, algorithms. Here you begin with one prompt that does the whole job. Input goes in, the model works out what to do, output comes out. On day one the application is almost entirely probabilistic.

Note what the model is doing here. It isn’t helping someone write the software; at first, it is the software. That’s the difference from AI-assisted coding, where the model’s output is code that a developer then owns.

Then you run it and watch where it goes wrong. It gets arithmetic wrong. It works through the same reasoning on every request. A common operation takes ten seconds when it should take ten milliseconds.

Each of those is a candidate to pull out into deterministic code: a function, a validation rule, a lookup table, a state machine, a tool the model can call. The model keeps whatever hasn’t been pulled out yet.

Keep doing this and the code grows while the prompt shrinks. The model may end up as a small part of the system, or go away entirely. That isn’t the aim, though. The aim is to put the line in the right place.

1 · First version

LLM

2 · Some behaviour extracted

LLM code + tools

3 · Mostly code

LLM code + tools

4 · Possible end state

code + tools
How the share of the system shifts over time. The proportions are illustrative, not measured, and stage 4 is optional.

The loop

The usual cycle looks roughly like this:

specify → design → implement → test → refine

This one looks like this:

specify → run → observe → determinize → repeat

The design step moves. You don’t have to understand the whole domain before you write anything. You start with something that works, badly, and you formalize the parts that turn out to need reliability, speed or predictability. Running the system tells you where those are, instead of guessing up front.

  1. Specify the behaviour you want
  2. Run it: the LLM plus whatever code exists so far
  3. Observe the results, failures and costs
  4. Is this piece worth turning into code?
    yesWrite it as code and verify it against real correctness criteria
    noLeave it with the model for now
Either way, go back to step 2 and run again.

Observing needs more than reading an output now and then. I’d log every model call: the input, the output, the cost, how long it took, and whether a person corrected the result afterwards. Coming from event sourcing, I see that log as an event stream. Questions like “which suppliers get corrected most?” or “what does this step cost per month?” are projections over it, and their answers tell you what to extract next.

An example: invoices to vouchers

Take bookkeeping. For every invoice you need a journal entry voucher, meaning a debit account, a credit account, a VAT code and amounts that add up.

The first version is one prompt. Give it the PDF and the chart of accounts, and ask for a voucher back:

Invoice  Cloud hosting, Germany
         EUR 60.00, no VAT charged
Rate     11.70 NOK/EUR
Debit    6553 Software     702.00
Credit   2400 Payables     702.00
VAT      Reverse charge, 25%
         175.50 in / 175.50 out

That works surprisingly often. Run it on a few months of invoices, though, and the problems are easy to list. It gets the VAT split off by a few cents. It books the same hosting bill to a different account in March than it did in February. Once in a while it uses an account number that doesn’t exist, or guesses the exchange rate.

I’d fix those one at a time. The exchange rate is looked up from the published rate for the invoice date; a model should never be guessing it. The VAT calculation is just arithmetic, so it goes into a function. The model shouldn’t be making up account numbers, so it gets the chart of accounts as a fixed list and anything else is rejected. A check makes sure debits and credits match before a voucher is saved. And most suppliers are the same every month, so they go into a table with their account and VAT code. If the supplier is in the table, the model isn’t called.

After that the model only sees what the table doesn’t know about, mostly new suppliers and the odd receipt. Once someone confirms how a new supplier should be booked, it goes into the table, and next month that supplier doesn’t go to the model either.

Tests are what make it safe

None of this works without tests. Every extraction is a change to behaviour that used to work most of the time, and you need to know it now works all of the time, and that you haven’t broken anything next to it.

It’s also worth being clear that deterministic doesn’t mean correct. A function that books an invoice to the wrong account will book it to the wrong account every time. Moving something out of the prompt makes it predictable. Whether it’s right still depends on the tests and on the domain knowledge behind them.

I like behaviour-style tests for this, in the BDD sense: scenarios written as given/when/then, about what the system should do rather than how:

Scenario: Remotely delivered service from abroad, no VAT charged
  Given we are registered for VAT in Norway
  And an invoice for software or hosting from a supplier outside Norway
  And the invoice charges no VAT
  When it is booked
  Then the expense is debited to 6553 Software
  And the VAT code is reverse charge, 25%
  And debits equal credits

I treat these scenarios as the contract for the feature, and they get used in two places. They’re the tests, of course. But they also go into the prompt, so the model doesn’t have to guess how to book a foreign software invoice without VAT. It’s told, in the same words the test will check. The same text ends up as the spec, the test and part of the prompt.

That keeps working as the system changes. The same scenarios run against the day-one prompt and against the code that replaces it. Once a scenario is handled by code you can drop it from the prompt, and the test still holds the code to it.

Running the scenarios also tells you what to extract next. Run them twenty times against the prompt and see which ones fail now and then. Those are the candidates. Once something is code, it has to pass every time.

It doesn’t have to be Gherkin. Plain unit tests, or a table of inputs and expected outputs, work too, though readable scenarios are easier to put in a prompt. The important part is that the expected answers come from someone who knows the domain, an accountant or the tax rules, and not from what the model produced before. The model’s old answers include its mistakes. They aren’t a spec.

Code is an investment

Deterministic code isn’t free. Someone has to design it, write it, test it and maintain it. So each extraction should pay for itself: fewer errors, lower running costs, faster responses, or guarantees a model can’t give you.

Plenty of behaviour won’t clear that bar. If something is rare, messy, ambiguous or keeps changing, it can stay in the prompt indefinitely, and that’s fine.

Why I think it matters

You could read this as a prototyping trick: hack it together with a prompt, then rewrite it properly. I think it’s more than that. It changes when you need to understand the problem. Instead of designing the whole solution before building it, you build a rough version and let real use show you where the engineering effort should go. The architecture comes out of what you observe, not out of a design document written before anyone used the thing.

It isn’t an architecture either. At any given point the system may look like any other LLM with tools. What’s different is how it got there, and how you decide what moves next.

It also gives you a number you don’t usually have. Whatever still goes to the model is the part of the domain you haven’t modelled yet, or decided not to.

Start with intent. Add determinism where experience shows you need it.