From Data Analytics to AI: The Technology Behind Modern Fraud Detection

A supplier changes its bank details at 3:12 p.m. Three invoices arrive before 5:00, each just below the level requiring senior approval. That evening, an unfamiliar device accesses the payment account. None of those events proves fraud. Together, they form the kind of pattern modern detection systems are built to recognise.

Fraud detection now combines transaction analytics, rule engines, machine learning, graph models, document intelligence and real-time controls. The goal is not to let software declare guilt. It is to identify the few events that deserve examination, preserve the evidence and give investigators a defensible starting point.

Fraud Starts as Data

Every digital business process produces evidence. Enterprise systems record purchase orders, invoices, approvals and payment releases. Identity platforms record logins and devices, along with permission changes. Document systems preserve contract revisions, while expense tools store receipts and reimbursement histories.

The important change is not simply that companies collect more information. Modern fraud platforms can connect records once reviewed in isolation. A suspicious invoice can be linked to the person who created the supplier, the approver who released the payment, the device used to change the bank account and earlier complaints involving the same vendor.

That connected view matters because fraud rarely appears as one obvious transaction. The ACFE’s 2026 study examined 2,402 occupational fraud cases across 143 countries. Those cases caused more than $3.4 billion in losses, and the median scheme continued for 12 months before detection. Fraud found within six months had a median loss of $40,000, while schemes lasting more than five years produced median losses above $1.1 million.

Raw records do not identify risk on their own. The next layer establishes what normal activity looks like.

Analytics Defines Normal

Data analytics calculates expected invoice values, ordinary payment times, usual login locations, standard approval paths and typical supplier behaviour.

Consider a logistics company that pays a carrier twice per month to the same account. If the carrier suddenly changes its payment destination, submits four urgent invoices and receives approval from a manager who has not handled that supplier before, the deviation becomes measurable. Analytics does not label the activity fraudulent. It shows how far it sits from the historical baseline.

Three forms of analysis usually work together:

  • Descriptive analysis identifies what changed, such as a rise in refunds or repeated invoice amounts.
  • Diagnostic analysis looks for the reason by connecting the change with supplier updates, staffing movements or device activity.
  • Predictive analysis estimates risk by comparing the pattern with earlier confirmed cases.

A useful baseline must be specific. Comparing every department against one company-wide average produces weak alerts because operating patterns differ by role, geography and transaction type.

Once normal behaviour is defined, rule engines can stop known schemes immediately.

Rules Remain Essential

Rule-based detection is older than machine learning, but it remains one of the most reliable layers in the stack. A rule can flag a duplicate invoice, hold a payment after a bank-account change or require a second approver when a transaction approaches a financial limit.

Rules are valuable because investigators can explain them. They know which condition was met and which record triggered the alert. The resulting action is visible.

Their weakness is rigidity. A fraudster can vary invoice amounts or rotate accounts. A large payment can also be split into smaller transactions that remain outside the threshold. Static rules produce additional noise when business conditions change.

Detection layer

Best at finding

Main limitation

Rule engines

Known schemes with clear conditions

Predictable logic can be avoided

Statistical analytics

Deviations from historical behaviour

Complex networks may remain hidden

Machine learning

Risk patterns across many variables

Scores can be difficult to explain

Graph analytics

Collusion and connected entities

Poor identity matching creates false links

Language AI

Inconsistencies across documents

Context can be misunderstood

The limits of fixed rules explain why organisations add machine learning rather than replace every existing control.

Machine Learning Ranks Suspicion

Machine learning evaluates combinations that are difficult to express as individual rules. A model may consider payment value, timing, supplier history, device identity, approval route and employee role at the same time.

Supervised models learn from cases investigators have labelled as legitimate or fraudulent. They can recognise combinations associated with account takeover and duplicate billing, along with expense manipulation. Their quality depends on those labels. Incomplete or inconsistent investigations become part of the training data.

Unsupervised models identify behaviour outside the normal pattern of an account or employee. They can also expose unusual supplier activity. This helps detect emerging schemes, but it creates more false positives. An emergency purchase or new international project may also look abnormal.

The correct output is a risk score, not a verdict. A model should reduce millions of events to a reviewable set and show which factors increased the score.

Adoption is rising but remains limited. In the ACFE and SAS 2026 benchmarking study, 25 percent of surveyed organisations used AI or machine learning in anti-fraud work, up from 18 percent in 2024. Another 28 percent expected to adopt it by 2028. Yet 82 percent considered explainability important, while only 6 percent felt completely confident explaining how their models reached anti-fraud decisions.

Machine learning finds unusual behaviour. Graph technology shows whether it belongs to a wider network.

Graph AI Exposes Networks

Many schemes look ordinary when transactions are viewed one at a time. Graph analytics changes the unit of analysis from the transaction to the relationship.

Employees, suppliers, bank accounts, devices, addresses and payments become connected entities. The platform may reveal that five suppliers use the same account, several claimants submit records from one device or a manager repeatedly approves payments to businesses linked through directors and phone numbers.

This is useful for collusion and synthetic identities. It can also expose layered payment networks because the network may carry a risk signal that no single invoice contains.

Graph results still require careful identity resolution. Two companies may share a registered office without any fraudulent connection. Investigators need confidence levels and access to the source records behind each inferred link.

Financial data and relationships capture only part of the evidence. Fraud also leaves clues inside language.

Language AI Reads the Record

Contracts, claim forms, emails, complaints and investigation notes contain information that transaction models cannot process effectively at scale. Language AI can extract entities and compare statements. It can also organise large document sets.

A procurement system can compare invoiced work with contractual deliverables. A healthcare review team can identify differences between billing descriptions and supporting documentation. A bank can group complaints that describe the same scam in different words.

Generative AI can build timelines and summarise repetitive files. It can also identify missing documents. Among organisations already using generative AI for anti-fraud work, the ACFE study found common uses in phishing and scam detection. Risk assessment and report writing were also common. Only 16 percent of respondents were using generative AI, although 58 percent planned future adoption.

The model should never become the evidence. Every material statement in a generated summary must link back to the invoice, message, contract or claim that supports it.

When transaction and network signals are combined with language analysis, detection can move from retrospective review to active intervention.

Detection Moves in Real Time

A real-time platform evaluates risk before the process is complete. If supplier bank details change shortly before a large payment, the system can check the new account against linked entities, review the approver’s history and compare the invoice with the contract.

The response should reflect the strength of the evidence:

  • Low-risk deviations can be logged without interrupting normal operations.
  • Medium-risk events can require supporting documents or an independent approver.
  • High-risk transactions can be paused while an investigator verifies the account and approval history.

This tiered structure prevents weak alerts from blocking legitimate business. A control that stops too many valid transactions will eventually be bypassed or ignored.

The platform should also record who reviewed each alert, which evidence was examined and why the case was cleared or escalated. Those decisions become both an audit trail and feedback for later model testing.

From Alert to Evidence

An AI alert begins an investigation. It does not finish one. Reviewers must return to original records and reconstruct the sequence of events.

A sound case file may include the invoice, supplier master-data changes, contract versions, approval history, login records, internal messages, model factors and earlier reports. The investigator should be able to explain why the alert appeared without relying on a black-box score.

Evidence integrity is equally important. Shared administrator accounts weaken attribution. Editable logs create uncertainty over timing. Exported spreadsheets may lose source metadata. Mature programs use named accounts, protected audit trails, documented model versions and recorded access to investigative material.

NIST’s AI Risk Management Framework organises AI oversight through four functions: Govern, Map, Measure and Manage. Applied to fraud systems, that means assigning ownership and defining the operating context. It also requires testing performance and responding to identified risks throughout the system’s lifecycle.

Most findings stay within internal compliance or security processes. A narrower issue arises when the detected activity concerns claims made to government programs.

When Public Funds Are Involved

Fraud analytics may uncover irregularities involving government contracts, healthcare reimbursement, research grants or other publicly funded programs. Payment records, approval histories, internal messages and repeated alerts may help distinguish a correctable billing error from a wider pattern.

When the suspected conduct involves knowingly false claims submitted for government payment, a person with direct knowledge may consult a False Claims Act lawyer to understand how the available records relate to the situation. This can become relevant when an internal review has stalled or evidence may be altered. Further reporting may also raise concerns about retaliation. The False Claims Act addresses knowingly false claims to the government and false records material to those claims.

The technology must still withstand scrutiny. A serious allegation cannot rest on a flawed model or incomplete data.

Where AI Detection Breaks

Poor data quality is the first failure point. Duplicate identities and missing fields create unreliable baselines. Inconsistent timestamps compound the problem. A sophisticated model trained on weak records produces precise-looking errors.

False positives create a second problem. If analysts receive thousands of low-quality alerts, high-risk cases disappear inside the queue. Performance should be measured against final investigation outcomes, not only test-set accuracy.

Fraudsters also adapt. They split payments, alter document wording, rotate devices and test controls until they understand the thresholds. Detection models require regular review against new behaviour.

Governance has not kept pace with adoption. The 2026 Techraisal study found that only 18 percent of respondents said their organisations tested AI models for bias or fairness, even though 86 percent considered accuracy important. Only 7 percent felt their organisations were more than moderately prepared to prevent or detect AI-fuelled fraud.

A defensible operating model should require four controls:

  • Investigators compare alerts with final case outcomes to track missed fraud and false positives.
  • Technical teams document data-source and model changes, along with permission updates, before production deployment.
  • High-impact findings require human approval and direct access to supporting records.
  • Independent reviewers test whether the system applies comparable standards across departments and locations.

These controls determine whether fraud technology produces reliable findings or automated suspicion.

The Next Detection Stack

The next generation of platforms will combine streaming transaction data, graph relationships, document analysis and investigator feedback in one case environment. Agentic AI may collect related records and build timelines. It may also identify missing documents, but people should approve decisions that freeze funds or accuse an employee. Any step that triggers external reporting should also require human approval.

Data lineage will become as important as model accuracy. Investigators need to know where each field originated and who changed it. They also need to know which model version used the field. Privacy-preserving methods will matter when companies compare sensitive records across departments or institutions.

The strongest platform will not be the one with the largest language model. It will be the one that connects reliable data, selects the correct analytical method, explains its alerts and preserves evidence in a form another reviewer can verify.

The Verdict

Modern fraud detection is a layered engineering system. Rules stop familiar schemes. Analytics establishes normal behaviour. Machine learning ranks anomalies. Graph AI reveals hidden networks. Language models organise documentary evidence, while real-time controls intervene before a transaction is completed.

AI provides speed and scale, but it does not establish intent. Credible detection still depends on clean data, explainable results, protected records and investigators willing to challenge the model. Technology can reveal connections that manual review would miss. Disciplined evidence handling turns those connections into findings that can withstand scrutiny.