Building HIPAA-compliant AI in healthcare software
The most common question we got in 2025 was "can we use OpenAI in
our healthcare product?" In 2026, that question has evolved.
Founders aren't asking whether to use AI anymore. They're asking
which architecture lets them ship LLM features without breaking
their compliance posture, slowing their roadmap, or building
something that can't survive an audit.
The answer depends on three architectural choices, and we'll walk
through them as we'd walk through them with you on a call.
1. The PHI boundary on the model
Where does the model run, and what does it see? Options range from
"nothing PHI ever reaches the model" (you strip identifiers before
inference) to "the model handles PHI directly under a BAA"
(OpenAI's Enterprise tier, AWS Bedrock with the appropriate BAA,
Anthropic's Claude via authorized providers). Each option has
implications for what features you can build, what data leaves
your environment, and what changes if you switch model vendors.
For most products, the right answer is a hybrid: structured
extraction can often run on de-identified inputs, while
conversational features that require context need the model
inside the BAA envelope. Our default architecture isolates AI
inference into a dedicated service with a clear input/output
contract, so you can swap the underlying provider without
rewriting product code.
2. Audit and observability for AI
Every AI inference that touches PHI needs the same audit trail as
any other PHI access. That sounds obvious, but most early AI
implementations skip it because the foundational logging
assumptions don't apply cleanly to model calls. We design AI
audit logging upfront: which user invoked the model, what
prompt was sent (or a hash if the prompt itself contains PHI),
what response came back, how long it took, what version of the
model.
The same audit log doubles as the dataset for evaluating model
quality drift over time. You don't want to discover that your
AI assistant got 12% worse after a vendor's silent model
update; you want to detect it from your own logs.
3. Human-in-the-loop where it matters
The fastest way to break trust with clinicians is to ship an AI
feature that bypasses their judgment on decisions they're
professionally responsible for. We design AI features as
assistive by default: the model proposes, a human disposes,
and the human's decision is the system of record. For some
workflows that's a UX detail. For others (clinical decision
support that meets the FDA's SaMD definition, for instance)
it's a regulatory requirement.
Tools we work with
OpenAI (via Azure OpenAI for BAA coverage), Anthropic Claude
(via AWS Bedrock with BAA), open-source models hosted on
dedicated infrastructure for the most sensitive workloads,
LangChain and the Model Context Protocol (MCP) for orchestration,
and the standard set of vector databases and embedding models.
We bias toward MCP as the integration layer between LLMs and
internal tooling. It's emerging as a useful standard for
composing AI features inside complex enterprise stacks.