- Legal and regulatory compliance is one of the highest-stakes places to deploy AI. A wrong or fabricated answer сan cost a real filing, a fine, or months of delay. Most AI tools sidestep this by staying vague or hedging everything.
- This case study describes a system built to do the opposite: make specific, checkable claims and back every one with a verifiable source.
The outcome is a multi-agent AI system that reviews documents like a law firm handling a complex case, with a specialist assigned to each legal area. It supports every conclusion with a verifiable quote from official legislation. We tested the system on real Ukrainian submissions covering eight regulated areas and 2,918 legal articles. In under two minutes, it completed a full multi-domain compliance check, identified genuine violations, correctly applied martial-law exceptions, and clearly flagged anything it could not verify rather than guessing.
The challenge: when regulatory complexity outpaces manual review
Professionals in regulated fields such as accounting, HR, legal services, land administration, public procurement, healthcare, education, occupational safety, and local government often file complex documents. These documents must comply with dense and frequently changing legal and regulatory frameworks. Even small mistakes can have serious consequences. A wrong checkbox, a mismatched date, an outdated form version, or a missing attachment can result in a rejected filing, a fine, or months of delay.
This problem is more challenging than it first appears because:
- Regulations move. In Ukraine specifically, martial-law statutes suspend individual articles of the Labour Code, Constitutional Court rulings strike down provisions mid-year, and entire codes get replaced (the Commercial Code was repealed outright in 2025). A document that was compliant last quarter may not be compliant today.
- Documents rarely sit inside one legal domain. A land privatisation application touches property law, martial-law transitional provisions, and local self-government authority at once. A construction contract touches labour law, accounting, and occupational safety at the same time. Checking it properly means evaluating it against several bodies of law in parallel, something a single generalist reviewer rarely has the time or full expertise to do well.
- Manual review doesn't scale. Expert review is slow and expensive, and there's no instant feedback loop to catch mistakes before submission.
Why generic AI can't be trusted with law
Generic AI can't be trusted with law, because simply pointing a chatbot at a document doesn't work in a regulated context. Generic LLMs hallucinate legal references — confidently citing article numbers that don't exist, or quoting laws in wording they never actually used. And even a citation that's verbatim correct can still be wrong: an article might be suspended under martial law, or struck down by a Constitutional Court ruling. That's a "meta-law" layer that keyword matching never sees.
Ukrainian adds another wrinkle: it's highly inflected, so naive keyword matching breaks on morphology, causing retrieval and document routing to fail silently, without any obvious sign that something went wrong. On top of that, without source attribution, a professional has no real way to verify whether an AI's feedback is grounded in an actual regulation or just made up.
Worse still, a false 'all clear' is worse than no answer at all. A tool that checks a document against the wrong domain's rules and then reports it compliant doesn't just fail to help — it actively misleads the user. Trustworthy validation requires verified sources, awareness of legal overrides, and calibrated honesty about what the system did and didn't check. That's a combination no static checklist or off-the-shelf conversational AI delivers.
The solution: a multi-agent approach to legal compliance review
The resulting system is a web application: a user uploads one or several documents (PDF or DOCX) and receives a structured, expert-level compliance report within minutes — findings highlighted in place, each with a verdict, a confidence score, a verified law quote, a suggested fix, and a link to the official source.
Rather than one model trying to know all of Ukrainian law, the architecture works like a law firm handling a complex case: it assembles a team of specialists.
- Parallel domain experts. The document is reviewed simultaneously by AI agents, each trained on curated material for one legal domain — general/civil law, labour and HR, accounting, public procurement, healthcare, local self-government, education, and occupational safety. Because domains run in parallel rather than sequentially, a document touching five domains takes roughly as long to review as one touching a single domain.
- Live web search for currency. Before issuing a verdict, a specialist can check a law's current wording in real time rather than relying solely on frozen training data — important in a legal environment that keeps shifting under martial law.
- Cross-check between specialists. A separate reviewing agent compares every specialist's conclusions against the others, looking for conflicts or gaps at the seams between domains. For example, it checks for cases where one rule appears to permit an action while a more specific rule forbids it.
- A verdict, a confidence score, and a citation on every line. Each finding is marked compliant, violation, or unverified, carries a confidence percentage, and points to the exact article or resolution behind it. No advice is given without a source to back it.
Under the hood: the technology behind the review
- A curated knowledge base, not scraped text. 30 laws (2,918 articles) were ingested from official zakon.rada.gov.ua print editions, plus 64 hand-curated, machine-checkable compliance norms, each anchored to specific articles. Every law quote is byte-verified against the ingested official text before it ships.
- Confidence-gated domain routing. A deterministic, lemma-aware Ukrainian morphology classifier routes each document to one of eight domains; an LLM refines only uncertain cases, and a low-confidence guess produces an honest "out of scope" report with suggested relevant laws rather than a false "compliant."
- Two extraction paths feeding one verification pipeline. Deterministic rule-based detection always runs — the system works with no API key at all — while an LLM extractor adds recall for norms that regex can't express. Both feed the same downstream verification gates, so LLM findings get no free pass.
- Cross-source corroboration and override detection. Each finding is checked against independent legal sources (two or more sources = corroborated); martial-law statutes that override a cited article flip the finding to a flagged conflict; Constitutional Court rulings that strike an article down remove it from corroboration entirely.
- An adversarial verifier. A second, independent LLM pass tries to disprove every violation before a user sees it. Refuted findings are discarded; uncertain ones are downgraded in confidence. During development, this refuter correctly rejected one of the system's own built-in norms as legally unsound (more below).
- Zero-hallucination link construction. The model never builds a citation or URL itself. Every link to a zakon.rada.gov.ua article anchor is constructed server-side from the verified knowledge base, and document quotes must anchor to the actual uploaded text.
- Freshness tracking. The knowledge base stores each law's official redaction date and compares it against the live legal register, so an amendment invalidates cached conclusions instead of silently going stale.
- Human-in-the-loop by design. Confidence is triple-encoded (colour, icon shape, and underline style, colourblind-safe) and kept separate from severity. The professional accepts or rejects every suggested fix inline, and only accepted fixes are written into the output document, exported as a corrected DOCX with genuine Word tracked changes.
A parallel deployment of the same pattern, built for accounting-form review (branded as e‑Expert), used semantic vector retrieval with multi-query RAG so each section of an uploaded form is matched independently against the knowledge base, plus a two-pass LLM-as-judge step that produces a quantified faithfulness score and flags any unsupported statements — surfaced to the user through an evidence-scoring dashboard (Evidence average, Evidence max, Faithfulness %).
Results: what happened when we ran it on real cases
Batch legal review. Seven real documents spanning three domains — a land privatisation application, several versions of a state social benefit declaration, and two simplified-taxation-system applications — were reviewed together in a single pass, producing 13 compliant findings, 11 violations, and 6 unverified items. The slowest single document in the batch finished in 96.6 seconds: a full multi-domain legal review in under two minutes.
Specific findings illustrated the level of nuance involved:
- A land application had requested the wrong permit type for its situation. In contrast, a separate martial-law exception it relied on for free land transfer was confirmed as correctly applied, citing the relevant transitional provision of the Land Code.
- A simplified-taxation filing left a required form-approval reference blank — itself grounds for rejection — and the system named both the current form identifier and the Ministry of Finance order behind it.
- A document citing the Labour Code's 120-hour annual overtime cap was flagged with the correct override: that article is currently suspended under martial law.
- The adversarial verifier caught a flaw in the system's own logic: an early version of a minimum-wage norm flagged any salary below the legal minimum as a violation, but since 2017 a base salary below minimum is lawful if a top-up brings total pay to the minimum. The refuter rejected the flawed test finding, and the norm was redesigned.
- Where evidence was genuinely insufficient — for instance, whether a military administration had assumed a local council's authority over a specific parcel — the system labelled the finding unverified rather than manufacturing a confident answer.
All eight of the client's professional domains (HR, accounting, law, occupational safety, education, healthcare, public procurement, and local self-government) were live and validated, covered by 62 backend tests plus seeded per-domain test documents, including a clean control document that stayed clean — confirming precision, not just recall.
Accounting-form review. Tested against a real Single Tax declaration for an individual entrepreneur, the system correctly identified that an annual-only appendix had been incorrectly attached to a quarterly declaration (severity 9/10), backed every finding with direct quotes from knowledge-base fragments and their sources, and explicitly listed items it could not evaluate due to insufficient data.
Value delivered: what professionals and organisations actually gain
- For professionals: structured feedback in minutes instead of days waiting on expert review; every finding independently verifiable against a cited source; a confidence and severity rating on every line to prioritise fixes; one review that covers a whole batch of related documents at once; and a corrected document with real, Word-compatible tracked changes rather than just a diagnosis.
- For organisations: fewer rejected filings and resubmissions; a single validation pass covering multiple domains simultaneously; reduced dependence on expensive specialist review for routine, high-volume document types; and a stronger compliance and security posture for regulated document workflows.
The strategic win: what makes this different from generic AI
The goal was not "another AI chatbot" but a demonstration of trustworthy artificial intelligence validation in a domain where a wrong or fabricated answer carries real financial and legal consequences:
- A team of specialists, not one generalist model — each legal domain gets its own deeply trained domain expert agent rather than one model spreading its attention across all of them.
- A source behind every claim — no fabricated statutes; every verdict traces to the exact article or resolution that supports it.
- Calibrated honesty as a product feature — instead of only scoring whether a document violates the law, we score the strength and completeness of the available evidence. We check whether our knowledge base actually covers the case and apply predefined passing thresholds. If the relevant legal sources are missing or insufficient, we say so clearly and recommend manual review.
- A domain-agnostic architecture — eight domains were shipped by repeating one proven template: ingest official texts, curate byte-verified norms, wire the same verification gates.
- Verification over generation — every legal claim passes through byte-verified quotes, cross-source corroboration, override detection, and an adversarial refuter before a user ever sees it. The model proposes; deterministic gates dispose.
Conclusion
This case study illustrates a broader point about where AI agents add value. They do not replace expert judgment but give it structure. Each domain can have a dedicated specialist. Every claim can be backed by a verifiable source. The system is also transparent about the limits of what it checked. Applied to legal document review, this combination turned a slow and expensive process into a compliance check completed in minutes.
The same pattern extends well beyond this one use case. Any workflow built around dense, frequently changing documentation, such as contracts, regulatory filings, internal policies, and technical specifications, can benefit from an AI system designed around verification rather than confident-sounding generation.
If you're exploring how AI could support your organisation's legal document workflows, or any other process that depends on rigorous control and verification, our AI consulting services can help you figure out what that could look like for your specific case.
FAQs
AI automates many tasks, including document review, legal research, contract analysis, compliance checks, due diligence, and risk identification. It lets lawyers process large amounts of legal information more quickly, so they can focus on more complex legal analysis and decision-making.
Certainly, AI can assist with searching and analysing legislation, regulations, case law, and other legal materials, thus making it easier to spot relevant provisions and possible legal issues. That said, you should verify AI results against official, authoritative sources, especially when accuracy and the current state of the law are essential.
Related Insights
Inconsistencies may occur.
The breadth of knowledge and understanding that ELEKS has within its walls allows us to leverage that expertise to make superior deliverables for our customers. When you work with ELEKS, you are working with the top 1% of the aptitude and engineering excellence of the whole country.
Right from the start, we really liked ELEKS’ commitment and engagement. They came to us with their best people to try to understand our context, our business idea, and developed the first prototype with us. They were very professional and very customer oriented. I think, without ELEKS it probably would not have been possible to have such a successful product in such a short period of time.
ELEKS has been involved in the development of a number of our consumer-facing websites and mobile applications that allow our customers to easily track their shipments, get the information they need as well as stay in touch with us. We’ve appreciated the level of ELEKS’ expertise, responsiveness and attention to details.