Skip to content

AI can investigate. Something still has to judge the evidence.

As AI takes on more engineering work, the systems judging production evidence may need to become more explicit. Why Ember is building deterministic engineering intelligence beneath optional AI.

5 min readEngineering LeadershipIncident ManagementSRE

AI is becoming part of the normal machinery of software engineering.

Coding agents can take an issue, change code, run tests and return with a pull request. Incident platforms increasingly describe AI agents that investigate across telemetry, code changes and previous incidents.

That is useful progress. It also creates a less obvious requirement. As more engineering work becomes probabilistic, the systems that decide what production evidence actually supports may need to become more explicit.

Not less.

Generation and judgment are different jobs

A good investigator forms hypotheses. A deployment happened shortly before latency increased. It touched the affected service. Similar changes have caused problems before.

That is enough to say: investigate this deployment.

It is not necessarily enough to say: this deployment caused the incident.

The distinction looks trivial when written down. During a live incident it is anything but. Several incomplete signals arrive together, and the first explanation that fits can quickly become the explanation, particularly when it is delivered clearly and confidently.

AI does not create this problem. Humans do exactly the same thing. But generative systems can produce coherent explanations much faster, across far more context, which makes a different set of questions more important than "what probably happened?":

  • What does each piece of evidence actually support, and what merely happened nearby?
  • What contradicts the current hypothesis?
  • What has already been tested, and what remains unknown?
  • If the same evidence were examined again, would the assessment be the same?

Intelligence does not have to mean an LLM

The word intelligence has become unusually easy to conflate with generative AI. They are not the same thing.

A system behaves intelligently when it maintains a model of what it knows and changes that model as evidence arrives. It can tell evidence that supports a claim from something merely associated with it. It can recognise when two observations are the same underlying signal rather than independent confirmation. It can preserve a contradiction instead of smoothing it into a cleaner story, and decline to make a claim when identity or causality is ambiguous. And it can remember why an assessment changed.

That is intelligence too, and it is the kind Ember is being built around.

Ember is an Engineering Intelligence Engine. The current runtime contains no AI or LLM integration. Its reasoning is deliberately deterministic: explicit causal support, independence rules, contradiction handling, fail-closed ambiguity and deterministic replay.

That is not an argument that an LLM should never sit inside Ember. It is an ordering decision: build the evidence substrate first.

An answer is not an evidence model

Two superficially similar architectures produce very different systems. The first is simple:

engineering systems
→ LLM
→ answer

Give a model logs, deployments, code, chat and telemetry, ask what happened, and receive a useful explanation. There is real value in that. But the answer can easily become the only durable artefact.

The second makes the evidence model the layer underneath:

engineering systems
→ deterministic engineering
  intelligence (Ember)
→ optional AI
→ humans and engineering tools

Here the durable object is the evidence model, and the generated explanation is a view over it. The evidence outlives any individual prompt, model or interface. A new model can narrate it tomorrow. An engineer can challenge it next week. A retrospective can replay it next month.

The intelligence survives the explanation.

Consider three deployments

14:03  deployment A
14:11  deployment B
14:18  deployment C
14:23  latency begins to rise

A changed the component showing the clearest symptom, so the team suspects it first. Any reasonable investigator, human or AI, would put it at the top of the list. A is rolled back. Nothing improves. That does not prove A irrelevant, but it weakens one of the strongest pieces of the hypothesis.

Then a symptom appears in a downstream dependency whose call path B changed. A customer-facing error follows on an independent path pointing the same way. C is investigated too, but the failing path never executed the code C changed.

The useful record is not simply B caused the incident. It is the state of the investigation:

  • A: initially plausible, weakened after rollback
  • B: independently supported
  • C: tested, currently unsupported

Each carries the evidence that moved it, and an observation five minutes later that contradicts B should move the record again. An Engineering Intelligence Engine preserves not only the winning explanation, but how the alternatives became stronger or weaker.

Deterministic does not mean certain

Determinism is not correctness. A deterministic system with the wrong rules will be wrong consistently. Production systems are messy, telemetry is incomplete and clocks disagree; no evidence engine should pretend incidents always reduce to certainty.

The value is narrower: given this evidence and these semantics, the assessment can be reproduced, its reasoning inspected and its rules challenged. And when the evidence cannot carry the claim, the system can abstain:

There are three plausible candidates. None currently has sufficient independent evidence to call causal.

A confident paragraph would be more satisfying. It would not be more useful.

The layer beneath the agents

None of this is an argument against AI. A deterministic layer underneath may make AI more useful. An engineer could ask why an assessment changed, what currently supports deployment B and what contradicts it, or for a summary written for the incident commander. Investigation, querying, synthesis and narration are excellent jobs for a language model.

The boundary is that narration should not silently become authority. An assistant can say two independent signals support a hypothesis; the evidence layer should be able to show which two.

That matters more as agents take on more of the engineering work. When one machine's conclusion becomes the input to another machine's action, provenance stops being a debugging nicety. You want to know what was actually observed, whether three signals were really independent, and which semantics produced the assessment at the time, with earlier states preserved rather than rewritten by hindsight.

Engineering organisations will not choose between deterministic software and AI. They will use both. The real question is which responsibilities belong to each.

Ember is being built as that quieter layer: one that knows where evidence came from, distinguishes a hypothesis from a supported claim, changes its assessment without forgetting why, and can say insufficient evidence rather than complete the story. See how it works.

As AI becomes more capable throughout the engineering stack, that distinction may matter more, not less.

Because AI can investigate.

Something still has to judge what the evidence actually supports.

Early access

Ember is being built on this thinking.

Evidence you can inspect, assessments that stay honest about uncertainty, and context that survives the incident.
Ember is in private development.