The Complete Guide to AI Fraud Investigation
A practical guide for fraud and compliance leaders

Contents
Most fraud teams generate more alerts than they can investigate. Detection keeps getting better, but working a case is still slow, manual work: evidence pulled from many systems, weighed against policy, and written up as a decision someone else can defend.
This guide is for fraud and compliance leaders deciding how much of that work an AI investigator should handle. By the end, you should know how the setup works, what a good investigation looks like, and how to roll one out safely.
Fraud investigations have become a strange mix of detective work and administrative work. Most begin in the transaction monitoring tool: an alert fires, and the investigator opens it to work out who the customer is and which transactions triggered it. The evidence lives everywhere else:
- The data warehouse, where assembling the transaction story can take several queries across several tables
- KYC and identity verification vendors
- Device intelligence and payment risk vendors
- Support tickets, chat logs, and the CRM
- The public web, where research is rarely one search and can turn into a rabbit hole
The hard part is connecting the pieces and deciding which signals actually matter.
This guide answers a practical question:
What can you hand to an AI investigator, and what should stay with your team?
Why now?
TL;DR: Detection systems create alerts, but investigators still have to assemble the evidence and decide what to do. AI can now handle much of that work, while fraudsters use the same technology to move faster.
Traditional fraud tools are good at producing alerts, scores, and signals. They're not built to carry out everything that comes next.
Fraudsters, meanwhile, are already using AI to test more targets and scale their operations. Generative AI could push US fraud losses to $40 billion by 2027, up from $12.3 billion in 2023, according to Deloitte's Center for Financial Services
Protecting good customers without adding unnecessary friction requires fraud teams to adopt better tools too.
That's where AI investigators come in. Across implementations, we've seen them help teams:
- Investigate KYC onboarding and transaction cases using the same underlying approach
- Combine internal data, risk vendor signals, and external research
- Conduct multi-step web investigation, navigating sites and extracting evidence
- Build an evidence-backed explanation for each disposition
- Resolve straightforward cases and escalate complex ones with the context already assembled
- Learn from prior investigations and validated reviewer feedback
The biggest change is that one investigation can now span structured and unstructured evidence: transactions and account records alongside documents, support conversations, websites, and images, all weighed against the team's policies and standard operating procedures (SOPs). A transaction rarely tells the whole story. Understanding it may require customer due diligence (CDD) or enhanced due diligence (EDD) on the customer behind it.
A typical investigation workload: Across Roe's implementations, flagged transaction investigations can take 30 minutes to hours when analysts gather information manually. They may pull from 8 to 10 systems, with 80-90% of the review time spent finding and assembling information. In comparison, an AI investigator can complete a case in 2-10 minutes, depending on investigation complexity.
There isn't one right operating model. Straightforward cases with clear evidence may be automated, while complex or high-impact cases stay human in the loop.
1. What AI fraud investigation means
TL;DR: Detection identifies activity that may be risky. Investigation explains what happened, weighs the available evidence, and determines what should happen next.
Stopping fraud before it happens is always better than investigating it afterward, and detection should catch what it can. But no team detects everything. New tactics have little history, legitimate customers can look suspicious for reasons a model hasn't seen, and broader rules buy more catches at the price of more false positives. Some of the most useful evidence never fits a model at all. It lives in KYC documents, support conversations, investigator notes, public websites, and other vendors' outputs, and it's different for every case.
Detection produces the alert, and that's where it stops. An AI investigator picks up from there: it gathers the relevant information, works out which signals matter, follows new leads, applies policy, and builds an evidence-backed disposition.
Fraud detection vs fraud investigation
| Fraud detection | Fraud investigation | |
|---|---|---|
| Main question | What activity might be risky? | What happened, and what should we do about it? |
| Typical output | Alert, rule hit, or risk score | Evidence-backed disposition |
| Data used | Signals selected in advance | Whatever relevant context the case requires |
| Workflow | Usually consistent across alerts | Changes as new information appears |
| Human effort | Creates the review queue | Resolves the cases in that queue |
A detection system might flag a payment because it came from a new device soon after a password reset and exceeded the customer's normal transaction size. Those signals still leave important questions:
- Was the device actually new, or did the customer replace a phone?
- Does the activity fit account takeover or legitimate behavior?
- Has the customer contacted support, or are related accounts involved?
The alert says something might be wrong. The investigation finds out what actually happened.
An AI investigator interprets signals in context
KYC providers verify identities and documents. Device intelligence vendors return data about devices, IP addresses, proxies, and account behavior. Payment systems flag unusual velocity, new counterparties, or activity outside the customer's expected profile.
Each vendor sees one slice of the case. One may return a score, another dozens of attributes, and a third a written explanation or document. An AI investigator can interpret those results together, then decide what to check next.
In practice, most teams pull these signals from several vendors at once: one for KYC, another for device intelligence, another for payment risk, each with its own format, portal, and login. Consolidating vendors doesn't fix the workload either. Even when one platform bundles several checks, alert volume can still outrun the team's capacity to investigate.
Interpreting those signals takes more than copying a vendor's risk label into the case. The investigator needs to understand what the vendor measured, how current the signal is, and how much weight it deserves alongside the rest of the evidence. A new device, for example, means something different when the customer recently contacted support about replacing a phone.
Consider a KYC review for a new business. The identity documents pass verification, but the company website was created recently. The business description doesn't match the products shown online, and the owner's address appears across several unrelated companies. The investigator can connect those facts, research the related entities, and explain whether the combination warrants enhanced review.
KYC and transaction investigations overlap
Teams often separate KYC onboarding from transaction investigations, but the work meets in the middle. A burst of high-value payments raises onboarding questions: what does the business sell, who owns it, and has the website or ownership changed? A thorough onboarding review, in turn, looks past the identity match to the business model, expected behavior, and online presence. Language models handle this overlap well because they work across tables, documents, and websites while carrying context from one step to the next.
The investigation can adapt without losing control
An AI investigator shouldn't follow the same fixed sequence for every case. A public records search may uncover a related business. Transaction analysis may reveal a counterparty that needs review. Some cases need three checks, while others require a deeper investigation.
That flexibility still needs boundaries. Teams decide which Data Connectors the investigator can use, what it can access, which policies it must follow, what evidence it must preserve, and when a person has to step in.
2. Decide where people stay in the loop
TL;DR: Use human review for complex, ambiguous, or high-impact investigations. Automate repeatable cases with clear evidence, and use a two-layer workflow to send exceptions to people.
An AI investigator can prepare a recommendation for a person or make the decision and complete the case. Choose the operating model by investigation type instead of applying one rule to every case.
Human in the loop
In a human-in-the-loop workflow, the AI investigator gathers customer and transaction data, interprets vendor signals, follows leads, checks the evidence against policy, and recommends a disposition. A person reviews the work before the decision takes effect.
Human review is a better fit when a case is long or complex, requires judgment that's hard to capture in policy, or carries a high cost if the decision is wrong. Missing information, conflicting evidence, and unfamiliar patterns are also good reasons to escalate.
The reviewer should receive the evidence, reasoning, citations, and proposed disposition together. Their job is to evaluate the judgment, not rebuild the case.
Can fraud investigations be fully automated?
In a fully automated workflow, the AI investigator completes the investigation and applies the disposition without routine approval. It may close a case, clear an alert, update case management, or trigger another action allowed by policy.
Full automation fits repeatable investigations with clear criteria, consistently available evidence, and a well-understood consequence if the decision is wrong. A person reviews exceptions instead of every case.
Getting here is not a day-one setting. It takes sustained iteration with the AI investigator: tuning the instructions, expanding the evals, and correcting cases until the error rate earns the autonomy. Most fraud teams aren't here yet, and that's fine. Full automation is something a workflow earns over time.

Two-layer alert triage handles the middle
A simpler first-layer investigator handles triage. It completes cases with clear evidence and familiar patterns, so false positives never reach the review queue. If the evidence is incomplete, contradictory, or unusual, it escalates with the work already assembled. The reviewer starts from a built case, not a raw alert.
People then focus on complex KYC reviews, longer transaction investigations, and cases with greater customer or financial impact.
Make the decision by case type
Set autonomy at the workflow or case-type level. A team might automate a narrow class of transaction alerts while keeping complex business onboarding reviews human in the loop. Even automated workflows should escalate policy exceptions and cases without enough evidence.
Ask four questions:
- How long and complex is the investigation?
- Are the decision criteria clear enough to apply consistently?
- Is reliable evidence available for every required check?
- Has the workflow been evaluated on enough representative cases?
3. Set up the AI investigator
TL;DR: Start with the procedures and completed cases you already have, then connect only the data needed for the first workflow. Keep permissions narrow and data warehouse access read-only.
Before an AI investigator can work a case, it needs to understand how the team investigates and where to find the evidence. Most of this setup can be reused as policies change.
Start with the SOPs you actually have
One thing we've learned across implementations is that very few fraud teams have complete, current documentation for every investigation. Procedures live across policy documents, checklists, training materials, case notes, and the experience of senior investigators.
Start with the material that exists, then compare it with how strong investigators handle real cases. AI can review SOPs and completed cases, identify recurring checks and decision criteria, and turn them into draft investigation instructions. Subject matter experts still validate the result, but they aren't starting from a blank page.
Those instructions should answer:
- What does the investigator need to check?
- Which sources should it trust for each type of evidence?
- What supports or contradicts each fraud hypothesis?
- When is the evidence sufficient?
- Which cases or policy exceptions require escalation?
Policies should set boundaries without forcing every case through the same sequence. Track changes to the instructions and test updates against past cases.
Separate hard rules from judgment calls
Not everything in the SOP should become AI judgment. Regulation-based requirements, internal thresholds, and fixed policies should stay deterministic. Encode them as hard rules in the detection stack or the workflow, where they run the same way every time. An AI investigator is probabilistic, and a mandatory check shouldn't depend on judgment.
Write the rest of the instructions so the investigator can be agentic with them. Instead of a fixed sequence of steps, describe what a strong investigator would weigh: which evidence matters, what supports or contradicts each hypothesis, and when the case is ready to close. The rules set the boundaries, and the investigator does the reasoning between them.
Add Data Connectors
Data Connectors link the AI investigator to the systems where evidence lives:
- Data warehouses like Snowflake, Databricks, and BigQuery, holding customer, account, transaction, login, dispute, and chargeback history
- CRMs and internal operations tools like Salesforce and HubSpot
- KYC, identity, device intelligence, and payment risk vendors like Socure, LexisNexis, and TransUnion
- Customer support platforms like Zendesk, Intercom, and Pylon, including email, chat logs, and call transcripts
- Case management systems and previous investigations
- Web portals and internal tools that may not offer an API
The investigator can query a data warehouse directly and call a risk vendor's API. When a system has no API, web investigation can navigate its interface and collect what a person would find by clicking through it.
Browser-based investigation is more than running a web search. The investigator may need to close pop-ups, handle cookie banners, click from a landing page to the relevant subpage, extract text or images, and preserve the source URL. The workflow also needs to tolerate routine website changes and stop when a login challenge or another control requires a person.
The investigator also needs to know that a device signal, support ticket, KYC record, and payment refer to the same case. Otherwise, adding systems only creates more disconnected facts.
Don't connect the entire fraud stack on day one. Start with the sources needed for one investigation type and add more when a case requires evidence the investigator can't reach.
Give the investigator only the access it needs, even as the use case expands. Data warehouse access should always remain read-only. Operational updates belong in other systems, with separate permissions and controls.

4. Run the investigation
With the setup in place, a case arrives and the investigation begins. It isn't a fixed checklist. The investigator chooses tools based on what it finds, revisits earlier assumptions when new evidence lands, and stops when the case supports a decision.

The lifecycle above is the map. What it looks like on a real case is easier to show than tell. The walkthrough below follows one illustrative transaction alert from intake to disposition, with the tool calls, findings, and sources along the way.

Two things are worth noticing. The investigator didn't run a fixed checklist: it researched the payee when the amount needed explaining, and went back to the warehouse a second time once the support history gave it new context to check. And every finding carries its source, so a reviewer can check any step without redoing the work.
5. Evidence, dispositions, and human review
TL;DR: Every important finding should link to source evidence. Reviewers need a concise answer first, with the ability to inspect the reasoning, citations, and investigation trace when necessary.
Decision accuracy isn't enough on its own. Fraud and compliance teams also need to understand why a decision was made and reproduce the work behind it.
Build an evidence-backed disposition
The AI investigator should preserve evidence throughout the case instead of reconstructing it at the end. Each important conclusion should connect to the record, document, query result, vendor response, or webpage that supports it.
A reviewer should be able to distinguish between:
- What the source says
- What the AI investigator observed
- How it interpreted that observation
- How the finding affected the disposition
Contradictory evidence belongs in the case record too. If one signal supports account takeover while another points to legitimate customer behavior, both matter. Hiding inconvenient evidence produces a weaker investigation.
The disposition should state the conclusion, which policy criteria were met, and what happens next. Depending on the workflow, that may mean clearing the alert, approving or declining an application, sending the customer a message or a request for information, changing the account status, placing the account on a watch list for continued monitoring, or escalating the case. The complete disposition, evidence, and citations should write back to the team's case management system, so the record lives where the team already works. For AML workflows, the same evidence package can seed the SAR narrative, with a person reviewing before anything is filed.

Make human review useful
The reviewer shouldn't have to redo the AI investigator's work. A summary template should define what the AI investigator presents and how the information is organized. The right level of detail varies by use case.
Most reviewers want the answer and main takeaway immediately. From there, they should be able to drill into supporting findings, source evidence, and individual investigation steps. The summary stays fast to scan without flattening a complex case.
Reviewers can approve the recommendation, correct it, request another step, or escalate. Record what changed and why so the feedback can improve testing and future investigations instead of disappearing into case notes.
Automated decisions still need oversight
Fully automated decisioning removes routine approval from individual cases. It doesn't remove human oversight from the program.
Teams still need to sample completed cases, review exceptions, watch for changes in data or behavior, and pause an automated workflow when results fall outside expectations. The same evidence standard should apply whether or not a person approves each case.
6. Give the AI investigator access to the right knowledge
An AI investigator can only reason over information it can reach. The model matters, but so do its data and knowledge sources.
Public information only goes so far
Web research provides company websites, business registries, marketplaces, and news. Some of the most useful fraud knowledge isn't public. It may live in prior investigations, analyst notes, internal typology documents, specialized risk vendors, private industry data, or consortium data.
For each investigation type, ask what experienced investigators know that the AI investigator can't find in the data warehouse or on the open web. That gap often points to another Data Connector or private knowledge source.
Build a living source of private fraud knowledge
A private knowledge base can store fraud tactics, validated examples, and investigative lessons that aren't available publicly. It should grow as new patterns appear and make relevant knowledge available when the AI investigator evaluates a case. It's also the idea behind Atlas, Roe's knowledge base of fraud and financial-crime typologies
Private knowledge still needs ownership and review. Teams should know where each entry came from, when it was last validated, and which workflows can use it.
Carry lessons forward with investigation memory
Investigation memory helps the AI investigator reuse context from prior work, including an earlier investigation of the same customer, a similar pattern, or validated reviewer feedback.
Not every past decision should become a permanent rule. Validate feedback before applying it more broadly and remove guidance that no longer reflects current policy.
Trusted investigation outcomes can also improve detection upstream. Once the AI investigator's dispositions hold up against review, recurring false-positive patterns become evidence for tuning detection rules, so the system produces fewer bad alerts while still catching the same fraud.

7. Quality, governance, and oversight
TL;DR: Evals should test the entire investigation, not only the final answer. Full traces, production audits, and change controls make performance measurable and give teams a way to intervene.
An eval is a repeatable test on a defined set of cases. The AI investigator completes each case, then the work is scored against standards set by experienced investigators. Run the same evals before launch, after a change, and throughout production.
Evaluate the investigation, not only the answer
A single accuracy score doesn't tell you enough. An AI investigator can reach the right disposition with weak evidence, skip a required policy check, or take twenty steps to complete a case that should take five.
Depending on the workflow, evals should measure:
- Whether the disposition and escalation decision are correct
- Whether the investigator found the right evidence and traced it to the source
- Whether the reasoning follows from the evidence and policy
- Whether it used its tools and time efficiently
- Whether it produced the required summary and case management updates
The shortest investigation isn't always the best one. The goal is to collect enough evidence without repeating work or following leads that won't affect the case.
Eval sets should include straightforward cases, difficult investigations, unusual fraud patterns, different customer types, and incomplete data. A workflow may be ready to assist a reviewer before it's ready for automated decisioning.
Make every investigation auditable
Evals measure performance across cases. Audits explain what happened in one production investigation.
A full agent trace should show the tools called, sources accessed, evidence collected, reasoning, and final disposition. Fraud teams can use it for internal review and produce a clearer record for sponsor banks, regulators, and other third parties.
Guard against fabricated evidence
Language models can state something confidently that the evidence doesn't support. The system should make that failure rare and visible. Require a citation for every material claim, flag any statement that lacks a source, and test for unsupported claims in the evals. Reviewers should be able to open the cited source next to the claim, so a fabricated detail has nowhere to hide.
Decide how long data is retained
Retention should be configurable. Keeping investigation history makes memory more useful, but the team controls the tradeoff: some run with full history, while others choose zero data retention or a short erase window, such as seven days. Whatever the setting, it should be documented and consistent with the team's data policies.
Put governance around changes
Governance defines who owns the evals, what performance is acceptable, and who can approve changes. Track the instructions, policies, summary templates, and Data Connectors used by each AI investigator, then rerun the relevant evals whenever one changes.
Set thresholds for increasing human review, pausing automated decisions, or returning to an earlier version when performance moves outside the approved range.
These controls also line up with what model risk teams and regulators expect from AI systems, in the spirit of SR 11-7 and the NIST AI Risk Management Framework: documented changes, measured performance, and the ability to explain any decision after the fact.

8. Decide what to build and what to buy
TL;DR: Internal builds fit narrow, stable workflows when the team can maintain every supporting layer. Buying can shorten the path to production, but teams still need to evaluate customization, data handling, auditability, and long-term dependence.
Building a prototype AI investigator has become much easier. Give a model a policy document, connect it to a few tools, and it may handle a carefully selected case well enough to look convincing.
Production is where the work expands. The investigator has to perform across messy cases, changing data, incomplete documentation, difficult websites, policy exceptions, and new fraud patterns. It also needs security, evidence handling, evals, monitoring, and operational controls.
An internal build may fit a narrow workflow with accessible data and support from engineering, AI, fraud, compliance, and security teams. It can also make sense when the capability is highly specific to the business and close to deterministic.
The ongoing work matters more than the first prototype. An internal team has to maintain Data Connectors, permissions, browser investigation, evidence, case summaries, evals, monitoring, and case management updates. Vendor APIs, websites, internal fields, policies, and fraud patterns will keep changing.
Buying can shorten the path to production when a team needs several Data Connectors, adaptive investigations, web research, audit trails, evals, and ongoing maintenance. It may also provide specialized fraud knowledge developed outside the organization's cases.
The tradeoff is dependence on an outside platform. Teams should understand how it handles data, how much the investigator can be customized, how performance is evaluated, and what happens when the models or product change.
The practical question isn't whether the organization owns everything or nothing. It's which layers the team needs to control directly and which ones it is prepared to maintain over time.
Before choosing, ask:
- How specific and stable is the workflow?
- Which data, policies, and actions require direct control?
- Can the team maintain connectors, browser automation, evals, and monitoring?
- What private fraud knowledge is available through each approach?
- How easily can the organization inspect, pause, or replace the system?
9. Roll out AI fraud investigation in stages
TL;DR: Start with one measurable workflow, test it on historical and live cases, and increase autonomy only after the evidence supports it. Shadow mode creates a safe bridge between offline testing and production use.
Start with one investigation type, either a KYC onboarding review or a transaction alert. Look for meaningful volume, a clear disposition, accessible evidence, and historical cases for evals.

Establish the baseline
Measure the current workflow using the metrics that fit the problem:
- Time to disposition and investigator time per case
- Case volume, backlog, and cost per investigation
- Decision quality, evidence completeness, and escalation rates
Test on historical cases
Build an eval set from representative investigations, including ordinary cases, exceptions, incomplete data, and the genuinely hard ones. Score the disposition, evidence, reasoning, policy adherence, and efficiency.
Run in shadow mode
In shadow mode, the AI investigator handles live cases alongside the existing process, but its recommendation doesn't affect the decision. The team compares its work with the human investigator's and uses the full trace to understand differences.
This catches changes in live data, unfamiliar cases, and Data Connector issues without changing customer outcomes.
Add human review
Reviewers can then use the AI investigator's summary, evidence, reasoning, and proposed disposition in live cases, inside the case management workflow they already use. Record what they approve or change, and feed the validated lessons into the workflow, the evals, and the investigation memory. That feedback loop is what makes the AI investigator improve with every stage.
Automate the cases that earn it
Expand autonomy by case type, completing straightforward cases that meet approved criteria while exceptions still go to a person.
Before expanding, confirm that the Data Connectors, policies, evals, review capacity, and pause controls are ready.
FAQs
Do AI investigators replace fraud analysts?
No. The AI investigator completes routine cases and assembles the evidence for the rest, and people handle the judgment calls, exceptions, and oversight. In practice teams shift from working queues to reviewing complex cases and tuning the system.
Can fraud investigations be fully automated?
Some case types can, eventually. Full automation fits repeatable investigations with clear criteria and reliable evidence, and it's earned through iteration and evals rather than switched on. Exceptions and policy edge cases still escalate to a person.
How long does an AI fraud investigation take?
Manual investigations of flagged transactions can take 30 minutes to hours across 8 to 10 systems. Across Roe's implementations, an AI investigator can complete a case in 2-10 minutes, depending on investigation complexity.
What's the difference between fraud detection and fraud investigation?
Detection flags activity that might be risky and produces an alert or score. Investigation explains what actually happened: it gathers evidence across systems, weighs it against policy, and ends in a documented disposition.
How does an AI investigator handle fraud patterns it hasn't seen?
The investigation is adaptive: unfamiliar evidence triggers deeper research, and cases that don't fit known patterns escalate to a person rather than being forced into one. Private fraud knowledge and investigation memory then carry the new pattern into future cases.
What do regulators expect from AI in fraud investigations?
Explainability and control: a full trace of how each decision was reached, measured performance through evals, documented changes, and the ability to pause or roll back. The same materials serve sponsor bank and auditor reviews.
Start with one investigation
Detection will keep improving, but investigations will remain necessary. Start with one workflow, test it on historical cases, run it in shadow mode, and let reviewers use the output before expanding autonomy.
Everything this guide describes is real work: Data Connectors to build and maintain, evidence standards to enforce, evals to write, traces to keep, memory to curate, and governance to run. You don't have to handle all of it on your own. We built those layers into Rori, Roe's Agent OS for fraud and AML investigations
Teams at Affirm, eBay, Kalshi, Dutchie, and Crossmint use Roe in production. Across live deployments, accuracy has started at 90%+ on day one, measured against analyst-graded ground truth, and the fastest team went from kickoff to a first production decision in one week.
If you're deciding how to start, reach out. We'll look at your caseload, show an investigation on cases like yours, and help you weigh build versus buy.
Get in touch
