01 / Research
Exploring AI-assisted incident intelligence
A feasibility study with Robert Gordon University
Examining whether AI and machine learning can help separate meaningful engineering signals from operational noise, and what such a system would need before anyone should act on what it concludes.
02 / The question
Can a machine tell an engineering signal from engineering noise?
What the study explored
- Whether AI and machine-learning techniques can separate meaningful engineering and incident signals from routine operational noise
- How incident evidence can be represented so that a model has something dependable to reason over
- What evaluation methodology is required before a result carries any weight
- How much of a model's reasoning can be surfaced, explained and challenged
What it did not attempt
- Independent validation of the Ember product, architecture or roadmap
- Autonomous diagnosis of production incidents
- Benchmarking against live customer environments
- Performance figures a study at this scale could not responsibly publish
03 / Why it matters
Incident intelligence is an evidence problem before it is a model problem.
Most engineering activity is not a problem
A running system produces a continuous stream of merges, deployments, alerts, metric movements and conversation. Almost all of it is ordinary. A system that treats every observation as a signal has not reduced the work of interpretation, it has only moved it.The meaning is in the connection, not the message
A single log line, alert or chat message rarely carries enough on its own. What makes an observation significant is usually what sits around it: the change that shipped, the service it touched, the window it fell in, and what was said at the time.A model is only as good as the evidence beneath it
Language models are fluent about incidents whether or not they have the evidence to be right. That is precisely the failure mode an engineering team cannot absorb, and it is why the study gave as much attention to evidence and evaluation as to modelling.
04 / Method
Scenario, evidence, analysis, evaluation, and back again.
01 / Scenario
A defined engineering situation with known ground truth, so a result can be checked against something.
02 / Evidence
The observations, signals and engineering records the scenario produced, captured with their provenance.
03 / Analysis
AI and machine-learning approaches applied to that evidence, producing classifications and scored outputs.
04 / Evaluation
Outputs measured against ground truth, with attention to how examples were grouped and separated.
05 / Learning
What the result changed about the method, the evidence, or the question — then the loop runs again.
The research evidence packs, and why they mattered
The study used structured incident scenarios and evidence generated through Ember's own engineering and test environment. Over the course of the project we built progressively richer, reproducible evidence packs for the university to work from.
Assembling them turned out to be one of the more instructive parts of the work. It forced a distinction we have kept since: what a scenario is defined to be, and what an AI system can legitimately infer from the evidence it has actually observed. Those are not the same thing, and a study that blurs them will flatter itself.
Evidence pack
- Scenario definitions
- Ground-truth information
- Ember-generated observations and signals
- Relevant engineering and incident evidence
- Model or scoring outputs
- Provenance needed to understand and reproduce the experiment
05 / Findings
Tractable — but requiring considerably more validation.
The work gave us evidence that the problem is tractable, while making clear how much more research and validation is required before anything here should be trusted in an operational setting.
AI and machine-learning approaches showed enough potential in this problem space to justify continued research and development. What the study did not show is that a system can autonomously diagnose production incidents, and nothing here should be read that way.
What meaningful performance was found to depend on
Training data quality and diversity
What a model learns from shapes the outcome more than the choice of model does. Narrow or repetitive data produces results that look better than they are.Rigorous evaluation methodology
Decided in advance and stated plainly, so that a result can be examined by someone who was not in the room when it was produced.Avoiding leakage between related examples
Closely related examples appearing on both sides of a split turn a score into a measure of memory rather than of ability.Appropriate grouping of evaluation datasets
How evaluation data is grouped changes what a number means. The grouping has to reflect how the underlying evidence is actually related.Explainability
Understanding why a model produced a result, not only that it produced one. An engineering team has to be able to disagree with it on the evidence.Reliable mapping between evidence and outcomes
Where the link between engineering evidence and its labelled outcome is loose, the label teaches the wrong lesson and the evaluation quietly rewards it.Validation against increasingly realistic scenarios
Structured scenarios are where this work has to begin. They are not where it can end, and each step toward operational conditions is a step that has to be earned.
06 / What it taught us
The most useful output was not a model.
Evidence before inference
An AI-generated conclusion is only as trustworthy as the engineering evidence underneath it. Improving the model is the easier half of the problem; improving what it reasons over is the half that decides whether the output can be relied on.
Reproducibility matters
Experiments and assessments need clear provenance, so that a result can be recreated, inspected and argued with. A conclusion nobody can retrace is not a finding, it is an assertion.
Incidents are contextual
Individual messages and telemetry points rarely contain enough information by themselves. Meaning emerges from connected evidence across systems and over time, which is a different problem from classifying a message in isolation.
Confidence needs boundaries
A useful system has to distinguish what it observed, what it inferred from that, and what remains genuinely uncertain. Collapsing those three into one confident sentence is how tools lose the trust of the people carrying the pager.
07 / Shaping Ember
Where the research is landing in the product.
Evidence provenance
Where an observation came from, carried alongside it.
Reproducible evaluation
Assessments that can be recreated and challenged later.
Evidence-gated claims
No conclusion the evidence underneath cannot carry.
Signal classification
Separating the meaningful from the routine.
Incident reasoning
Connecting evidence across systems and time.
Engineering context
The change, the service, the window, the conversation.
Model evaluation
Measuring the right thing, and saying how it was measured.
Explainability
Why a result, not only what the result was.
Learning and validation
Workflows that keep testing the system against reality.
Read alongside how Ember works, the throughline is the one the study kept returning to: a claim is only worth making if the evidence behind it can be produced on request.
08 / What is next
A useful first step, and a longer road.
Areas of interest for further research
- Larger and more diverse training datasets
- More rigorous grouped evaluation methodologies
- Advanced language-model approaches
- AI-assisted signal classification
- Incident prediction
- Adaptive risk assessment
- Operational decision support
- Evaluation against richer real-world engineering scenarios
09 / Thinking
The thinking this research sits next to
10 / Talk to us
