An alerting system that produces more alerts than anyone can investigate has not reduced risk, it has relocated it. The client's compliance team received a daily list of flagged transactions with no context, and spent most of its time assembling the same background information before it could judge anything. ELIVTECH built a platform where that assembly is done by agentic workflows before the alert arrives, so an analyst opens a case that already contains the evidence.
The client operates a compliance function monitoring transactions, counterparties and vendor relationships against regulatory and internal policy. Detection ran as overnight database queries producing a flat list of flags, and each one required an analyst to gather counterparty history, related transactions, sanctions and adverse media checks by hand before forming a view. Most of the working day was retrieval rather than judgement, and the backlog grew faster than it cleared. ELIVTECH rebuilt detection on an Elasticsearch data platform fed by Python pipelines, added agentic workflows that assemble the evidence behind each alert, and used generative models under strict grounding rules to draft narratives that cite the records they came from.
At a glance
The challenge
Detection that produced flags, and a team that had to build every investigation from nothing.
Alerts without context
Detection produced a transaction reference, a rule name and a date. Everything needed to judge whether the flag mattered had to be retrieved separately from several systems. The alert marked the beginning of the work rather than carrying any of it, so volume translated directly into hours.
Data spread across disconnected systems
Transactions, counterparty records, vendor master data, sanctions lists and prior case history lived in separate systems with different identifiers. Establishing that two records referred to the same counterparty was itself manual work, repeated on every investigation.
Overnight detection only
Rules ran as batch queries after the close of business, so anything occurring during the day surfaced the following morning at the earliest. For patterns that develop within a single session, the detection window was longer than the window in which action was still useful.
Rules that could not be tuned safely
Thresholds were embedded in query code with no way to test a change against historical data. Nobody could predict how many additional alerts a tightened rule would produce, so rules were left as they were and the team absorbed the consequences.
Case narratives written from scratch
Every closed case required a written rationale for the file. Analysts composed these individually, so quality and structure varied, and preparing a regulatory response meant reading through inconsistent free text rather than querying a structured record.
No measurement of what detection was worth
Outcomes were not linked back to the rules that raised them. The team could not say which rules produced useful alerts and which produced noise, so effort was distributed evenly across detection logic of very uneven value.
Our solution
Assemble the investigation before the analyst arrives, and ground every generated word in a retrievable record.
Case Analysis and Entity Resolution Design
Five weeks working through closed cases to establish what analysts actually retrieve before forming a judgement, which became the specification for the agent tools. Alongside it, an entity resolution model was designed to link counterparty records across systems with differing identifiers, since almost every investigation begins by establishing that separate records describe the same party.
Python Ingestion and Elasticsearch Risk Store
Python pipelines ingest transactions, counterparties, vendor data, sanctions lists and case history into Elasticsearch with resolved entity identifiers attached at write time. Detection runs continuously against the stream rather than overnight, and rules are held as configuration with a backtesting mode that replays a proposed change over historical data before it goes live, reporting the alert volume and the historical cases it would have caught or missed.
Agentic Evidence Gathering
When a rule fires, an agentic workflow runs a defined toolset over the risk store: counterparty history, related transactions, sanctions and adverse media status, prior cases and peer comparison. The agent decides which tools are relevant to that alert type and pursues follow-up questions, but can only call declared tools against indexed data, so its scope is bounded by design. Every tool call and result is logged with the alert, which means an investigation can be reconstructed later including what the agent looked at and did not find.
Grounded Narratives and Analyst Workflow
A generative model drafts the case narrative strictly from the gathered evidence, with every assertion linked to the record supporting it and no capacity to introduce facts the tools did not return. Analysts review, correct and decide, and their decisions feed rule performance reporting so detection logic can finally be assessed on the outcomes it produces. Where the evidence does not support a statement, the model is required to say so rather than fill the gap, and the analyst sees the absence explicitly.
Technology stack
Results
Measured once continuous detection and agentic evidence gathering were serving the full case load.
Before and after: platform engineering measures
- Alerts delivered as a reference, a rule name and a date
- Analysts retrieving background from several systems by hand
- Counterparty identity reconciled manually on every case
- Detection running overnight as batch queries
- Thresholds embedded in code with no way to test a change
- Case narratives composed individually in free text
- No link between detection rules and case outcomes
- Alerts arrive with the evidence already gathered
- Agent tools retrieve history, related activity and screening results
- Entity resolution applied at ingestion, not per investigation
- Detection running continuously against the ingestion stream
- Rules held as configuration and backtested before release
- Narratives drafted from evidence with every assertion cited
- Outcomes fed back into rule performance reporting
Project timeline
Case Analysis and Entity Modelling
Review of closed cases to specify agent tooling, entity resolution rule design across systems, data source and retention assessment, and the governance model for automated evidence gathering agreed with compliance.
Data Platform and Continuous Detection
Python ingestion pipelines across all sources, Elasticsearch risk store with resolved entity identifiers, streaming detection replacing overnight batch, rules as configuration, and the backtesting harness over historical data.
Agentic Evidence Gathering
Declared tool implementations over indexed data, agent orchestration with bounded scope and step limits, per-alert-type tool selection, full logging of every tool call and result, and evaluation against analyst-gathered evidence on historical cases.
Narrative Generation and Analyst Workflow
Grounded narrative drafting with per-assertion citation, refusal behaviour where evidence is absent, analyst review and correction interface, case decision capture, and rule performance reporting built from outcomes.
Assurance and Supervised Launch
Shadow running against live alerts with analyst comparison, model output review for grounding failures, monitoring on tool latency and agent step counts, and a supervised launch keeping analyst decision authority unchanged.
Key takeaways
What shaped the engineering decisions
- Automate retrieval, not judgement: Having a model decide whether an alert matters was rejected outright. The work that consumed the team was gathering evidence, and that is what the agent does, leaving the decision entirely with the analyst.
- Bound the agent with declared tools: An agent with open-ended access is impossible to assure. Restricting it to a declared toolset over indexed data, with step limits and full call logging, made its behaviour reviewable and its scope provable.
- Ground every generated assertion: A narrative that cannot be traced to records is unusable in a compliance file. Requiring each assertion to cite the record supporting it, and having the model decline where evidence is absent, made generated text something a regulator can follow.
- Resolve entities at ingestion: Reconciling counterparty identity per investigation repeated the same work on every case. Attaching resolved identifiers at write time made history retrieval a single query rather than an exercise.
- Backtest rules before releasing them: Thresholds in code could not be changed safely because nobody could predict the alert volume. Replaying a proposed rule over historical data made tuning an informed decision rather than an experiment on the live team.
- Close the loop from outcome to rule: Linking case decisions back to the rules that raised them let the team see which detection logic was producing value, which is the only basis on which effort can be allocated sensibly.
Want results like these?
Let's discuss how ELIVTECH can drive measurable outcomes for your business.
Start your project