16 September 2026/Terje Ennomäe
Coming to Leadstream: agent orchestration and Quality Score

This is a short pre-launch note. On 1 October 2026 we plan to release a new version of Leadstream. Dates can move, but that is what we are working towards — and we would rather tell you what is coming, and how we are building it, before the launch than after.
Two things are inside. One is an orchestration tool for our AI Agents. The other is a genuinely new capability: Automatic Quality Score, which evaluates customer service calls automatically, at scale. The second one is also a high-risk AI system under the EU AI Act, and that has shaped almost every design decision behind it. This post explains both, and is honest about the limits.
One place to run every Agent
Over the last two years Leadstream has grown a family of AI Agents — call summaries, automated topic detection, failure-demand spotting, repetitive-call analysis and others. Useful individually, but each with its own corner of the product. If you wanted to know what an Agent was actually doing, or change how it behaved, you had to go looking.
The orchestration tool brings all of them into one place. You can:
- Configure every Agent from a single view, rather than hunting through separate screens.
- Test an Agent on real conversations before you rely on its output.
- See what it does — the inputs it reads, the output it produces, and the exact configuration it ran under.
It is a small idea with a large effect: when Agents live in one place, they stop being magic and start being something you can inspect, trust and improve. That same foundation is what makes the next feature possible.
Automatic Quality Score
Most quality programmes review a tiny sample of conversations — often a handful of calls per agent each month. It satisfies a box; it does not tell you what is happening on the other 99% of calls. Automated evaluation changes the maths: you can score every conversation instead of guessing from a few.
Automatic Quality Score produces, for each call:
- One comparable score — a single figure you can track over time and across teams.
- 5 to 20 weighted quality dimensions — you decide what "good" means and how much each part counts.
- Compliance checks — did the required steps happen?
- Resolution labels — was the customer's issue actually resolved?
- Coaching text — a short, specific note a team lead can act on.
All of it runs automatically, for a fraction of the cost of manual listening. This is the same problem our quality assurance work has always addressed — just at full coverage rather than a sample.
Why we say "high-risk" out loud
We will say it directly: Quality Score is a high-risk AI system under the EU AI Act. AI that evaluates or monitors the performance and behaviour of employees falls under Annex III, point 4(b). When a score can inform coaching, reviews or decisions about people, the regulator does not treat it as casual analytics — and neither do we.
So we started from the design principles, not from the features. Building AI Act-ready workforce analytics means the system has to be trustworthy by construction, not by promise.
In plain terms: nothing happens in a black box. Every step is logged, so you can always see what an Agent did and why. Agents run on fixed revisions — a revision cannot be quietly changed, and any update becomes a new one, so every score points back to an exact Agent and version. Only people can touch prompts and configuration, and the people who build an Agent are not the ones who test it. Access is role-based and kept to the minimum needed. You test before anything goes live, you can see both what goes into the model and what comes out, and it is model-agnostic — you are not locked to a single vendor.
We do not think this part is a compliance cost. It is the reason you can trust a score at all.
The half of the work nobody demos
Here is the part that never shows up in a product demo. To put a high-risk system on the market, the software is only half of the job. The other half is written:
- a risk management system,
- a data governance description,
- technical documentation,
- logging and record-keeping,
- instructions for use for the deployer — purpose, limitations, human oversight,
- accuracy and robustness evidence,
- a quality management system,
- conformity assessment, a declaration of conformity, CE marking and EU database registration,
- post-market monitoring and incident reporting.
We are writing all of that in parallel with the code. It is slower. It also makes the product better, because you cannot document a limitation you have not first understood. Working through the documentation forces the honesty that the next section is made of.
The limitations, stated plainly
No score is the whole truth, so we will name where this one stops:
- The score is text-based. It reads the transcript, so tone is not included — a technically correct call delivered coldly can still read well.
- Prices are not verified against your price list. If an agent quotes a figure, Quality Score does not check it against your catalogue.
- AI judgement can vary. Two near-identical calls may land a point apart, so periodic calibration against human QA is advised.
None of these are reasons not to score every call. They are reasons to use the score as a fair, wide net that points humans at what matters — not as a verdict that replaces them.
Frequently asked questions
When is this available?
We are aiming for 1 October 2026. Dates can move, but that is the target, and we will share more detail as it firms up.
Does Automatic Quality Score replace human QA?
No. It gives you full coverage and a consistent score, then routes attention to the calls and coaching moments that matter. Humans still calibrate the system, own the prompts and make the decisions about people.
Why call it "high-risk" — isn't that off-putting?
It is simply what the EU AI Act says about AI that evaluates employee performance (Annex III, 4(b)). Naming it is the honest starting point, and the design principles that follow — auditability, fixed revisions, human oversight — are what make the score dependable.
Will my existing Agents change?
No. The orchestration tool gathers your current Agents into one place to configure and test; it does not change what they already do.
Where to go next
- The pillar guide: Automated call-centre quality assurance
- How a score is built: Inside an automated quality evaluation
- The regulation angle: The EU AI Act and employee performance assessment
- The product: Automatic quality assurance
Want an early look at agent orchestration and Automatic Quality Score on your own conversations? Book a demo and we will walk you through it before launch.