AI in fintech: speed without losing audit posture
The question we got in 2025 was "can we use OpenAI for our
fintech product?" In 2026, founders ask "how do I add AI features
to my fintech without breaking my SOC 2 or making my next audit
a nightmare?"
The answer comes down to three architectural choices.
1. The data boundary on the model
What data does the model see, where does inference run, what
gets logged? Options span from "nothing customer-data ever
reaches the model" (de-identify or tokenize before inference)
to "model runs inside the BAA / DPA envelope on Azure OpenAI,
AWS Bedrock, or dedicated infrastructure." Each option has
implications for which features you can build, which vendors
you're locked into, and what changes if you switch providers.
Our default architecture isolates AI inference into a dedicated
service with a clear input/output contract. The rest of the
application doesn't know it's calling an LLM: it calls a
structured service that happens to be model-backed today. This
means we can swap providers, run multiple in parallel for A/B
comparison, or fall back to deterministic logic when models
are unavailable.
2. Auditability of every AI decision
Regulators and auditors will eventually ask "why did this
AI-generated decision happen." Every LLM call that influences
a customer-facing outcome (fraud flag, credit decision, AML
alert, account closure) needs a complete audit record: input
(or its hash if sensitive), model version, prompt template
version, output, downstream effect, and the human who reviewed
it if applicable. This audit log is also your dataset for
measuring model quality drift over time.
3. Human-in-the-loop where stakes are high
A model that approves a $1,000 loan without human review is a
regulatory liability. A model that drafts a SAR for a compliance
officer to review and file is an efficiency multiplier. The
design decision is which workflows can run autonomously and
which require a human. We tend to bias toward assistive AI (the model proposes, a qualified human disposes) for any decision that touches a customer's money or eligibility.
What we've shipped
Audit Hero is the canonical example. It's a New York fintech
startup that automates the 401k audit process for CPAs and
Third-Party Administrators: work that previously took analysts
30 to 40 hours per engagement. We built the MVP: structured
extraction with LangChain and OpenAI, Finch API integration for
payroll data ingestion, anomaly detection that surfaces audit
risks (missing documents, mismatched figures, outlier
transactions) for analyst review. The architecture keeps the AI
behind a structured service contract, logs every inference, and
gates judgment-bearing decisions on human review, so the CPAs
and TPAs using the platform have confidence in workpapers
because the AI doesn't replace judgment; it makes judgment fast.
Tools we work with
OpenAI via Azure OpenAI (BAA / DPA coverage), Anthropic Claude
via AWS Bedrock or direct enterprise access, open-source models
hosted on dedicated infrastructure for the most sensitive
workloads, LangChain and the Model Context Protocol (MCP) for
orchestration of multi-step financial workflows. MCP in
particular is emerging as a useful pattern for composing AI
features inside complex regulated stacks: see our notes on
MCP in fintech AI.