Case studyApplied AIAll work

The model proposes. A person decides.

An internal AI platform for a lending-technology company: an assistant that can act across forty-seven tools, finance queues where the model proposes and a person confirms, and a browser agent whose limits are written in code rather than in the prompt.

Client A lending-technology company, North AmericaEngagement A team of three · 2026 — presentPractice Applied AI

The client sells AI decisioning to regulated lenders. Its own operations, like most companies', ran across a dozen tools that did not talk to each other: accounting, time tracking, tickets, documents, email, chat, source control. Every month-end, every client report and every release meant somebody stitching them together by hand.

We built the platform that replaces the stitching. It does three things. It unifies the company's operational tools behind one surface. It puts an AI assistant on top of that data that can answer questions, build reports and act in those tools. And it runs the finance and delivery work that needs judgement — cost allocation, budget variance, period close, release notes — as queues where the model proposes and a person decides.

It is used by the company's own staff: finance, operations, release management, delivery leads and leadership. It is not customer-facing, and it holds no borrower data by design. But because the company answers to regulated lenders, permissions, audit and human review are product features rather than compliance overhead.

What makes this hard

  • The assistant can act. An AI that can send an email, file a ticket or edit a calendar makes mistakes that leave the building. Being wrong is no longer a bad answer on a screen.
  • The numbers are money. Finance cannot use an assistant that might invent a figure, round a total or restate a transaction's accounting.
  • Permissions have to bind the assistant too. If a user may not see payroll, asking the assistant nicely cannot be a way round it.
  • Everything it reads can talk back. A web page, a document or an email can contain text written to instruct a model. The assistant reads all three.
  • Model calls have no natural ceiling. An agent in a loop, or a scheduled job over a backlog, can spend real money before anybody notices.
  • Agentic tasks are long. A multi-step run outlasts the load balancer's idle timeout, a page reload and a laptop going to sleep.

A prompt asks the model to behave. Code decides what actually runs.

Facts and decisions

Every finance feature rests on one distinction, and the rest of the design follows from it.

A fact

Read verbatim, never set by AI

Something a source system knows: an invoice total, an accounting class, a time entry. Facts are never invented, never inferred and never written by the model.

A decision

Proposed by AI, confirmed by a person

Something only the business knows: which client a cost served, which budget line it belongs to. The model proposes, shows its reasoning, and a person accepts or overrides — every time, with an audit trail.

This is why the platform is not "an AI that does the books". The model never touches the arithmetic.

An assistant that can act — and stops before it does

The assistant runs a tool loop across forty-seven built-in tools — email, calendar, documents, chat and project management — plus a query tool over the company's data warehouse, a report builder and a browser agent. Remote tool servers can be connected per user over the Model Context Protocol, and their tools join the same loop.

One turn that wants to send an email
1Read tools runSearching the inbox, reading a document, querying the warehouse. These execute at once, in parallel, under the same permission checks as everything else.
2A write tool is calledSend an email. The run pauses here.
3The person sees exactly what will happenA confirmation card listing each action and its parameters. Because the reads already ran, it shows a finished draft rather than placeholders.
4Approve, and only thenThe pending actions run and the loop continues. The pending state is single-use and expires, so an approval cannot be replayed.

Every write — send, post, create, update, delete, publish, accept an invitation — pauses the same way. Nothing leaves the building without a person saying so.

Two details make long agentic work reliable. The response streams, with a heartbeat every fifteen seconds, because the load balancer drops a connection that is idle for sixty — which is otherwise how a long run turns into an error page. And the run is detached from the request that started it: it keeps going whether or not anybody is watching, and a reload or dropped connection re-attaches to the same task and replays its progress.

The browser agent: autonomy with its limits in code

Some goals have no API. For those, the assistant starts a second, inner agent that operates the user's already signed-in browser tab — one action at a time, checking the result before the next.

Its limits are not instructions in its prompt. They are code, and the model cannot talk its way past them.

Allowed

A closed list of sites

Email, documents, calendar, chat, tickets and project tools, plus their sign-in pages. Anything else is refused — including acting on a tab that has wandered off the list.
Never

A hard deny list

The accounting system and time tracking. The browser may never operate them, not even to read. The deny list wins over everything.
Ask first

The point of no return

A click labelled send, submit, publish, pay, delete, share or approve — or Enter in a message box — pauses mid-run for the person's approval. A search box deliberately does not.

Around those sit step limits, a wall-clock limit, stall detection and a stop button, with every step narrated in plain language as it happens. And the prompt-injection posture is stated once and enforced everywhere: page content is data, never instructions.

In the finance queue, the model goes last

Staged accounting transactions arrive in a review queue with a proposal attached: which client, which budget line, how it should be treated, how confident, and plain-English reasoning that cites the rule it applied. A reviewer accepts, edits, rejects or asks for a second look.

The rebuilt pipeline asks, for each line, a chain of questions in order of how much each answer can be trusted — and stops at the first that answers.

The source system says soFact

The accounting system already names the customer. Written straight through; nothing proposes anything.

A ruleCertain

A vendor or class that always means one client. Deterministic, and stored as data rather than code, so finance can maintain it.

What a person decided beforeEvidence

The mapping a reviewer confirmed for this vendor last time. Every confirmation shortens the next month's queue.

The modelLast

A model reading the description — and only when nothing above has answered. Running it over a line a rule already answers adds a chance of disagreeing with a certainty, and there is no version of that trade worth making.

What the model sees is deliberately narrow: the vendor, the account, the class, the line description and the list of possible clients. No amounts, no dates, no totals. Not for secrecy — it is the company's own data — but because a model that can see an amount will use it, and "this is large, so it is probably the big client" is a correlation, not a reason.

What it may return is narrower still: a client, a confidence and a reason that must quote the text it relied on. It cannot express an account, a treatment, a split or a period. A proposal that could restate the accounting would, eventually, restate the accounting.

Finance tunes the rules themselves. The document that grounds the model lives in an in-app knowledge base, and editing it changes the next run without a deploy.

The assistant is bound by the same permissions

Access to the warehouse is default-deny in three layers: which source a role may read and at what fidelity, which rows, and which columns are masked. It is enforced on the generated SQL before it executes, and the dashboards and the assistant go through the identical chain. When masking applies, the model is told not to infer or reconstruct what it cannot see — and a refusal is relayed plainly rather than routed around.

Identity comes from the signed session only. A user id supplied by the browser — or by the model — is ignored.

One rule shows the thinking well. During the finance rebuild, two generations of warehouse tables exist side by side, and the assistant is made unable to read both at once. A prompt telling it to use only the new tables would not be enough: if the model inferred an old table's name, or read it in an error message, it could sum across both and produce a number roughly double and entirely plausible. So the other generation is denied at the SQL layer, for every role, including administrators.

Every call has a price

Every model call in the codebase goes through one gate that checks the user's daily and monthly caps before the call and records the tokens and cost after it, labelled by feature. Spend is attributable per person and per operation, and a failed call is ledgered too, so error rates stay visible.

The expensive paths are designed down. The system prompt and the tool definitions are cached, which on a multi-step run saves latency and cost on every turn after the first. Users choose a model tier per conversation, with small models doing the titling, summarising and memory work. And the unattended model in the finance pipeline is off by default: arming it is a decision, not a default, and when armed it runs under a per-cycle cap so a first-run backlog cannot spend the day's allowance and starve the person waiting on an interactive request.

Cutover against the ledger, not the old system

The rebuilt finance pipeline runs in parallel with the original. The gate for switching over compares it with the accounting system's own trial balance — never with the old pipeline, because a new system that matched the old one would faithfully reproduce the defects that prompted the rebuild. Coverage is measured by value rather than row count, and a night with no run is not a night that passed.

Release notes, with the guardrails in code

A release-management agent writes the notes for a product shipped to several lender clients on staggered schedules. It joins what source control says changed with what the ticketing system says shipped, flags the gaps either way, and produces three drafts in each client's voice: a publishable page, a client email and a chat summary.

The rules that matter are enforced on the final text, not merely requested in the prompt: never name an individual engineer, never invent a deploy date, mask anything shaped like a credential. Everything is marked as a draft for review, edits are shown against the AI's draft, and sending is gated — including a rule against Friday production releases whose override needs a written reason.

Honest status

The platform documents what is solid and what is not, and so does this case study.

AreaState
Assistant: tool loop, approvals, reports, memory, budgetsLive, in daily use
Browser agent with its allow and deny listsLive, in the desktop app
Warehouse sync from accounting, time tracking and ticketsLive — every fifteen minutes to weekly, by source
Release-notes agentLive
Three-layer permission enforcementBuilt, running in audit-only mode: every decision is evaluated and logged before it is switched on to deny
Rebuilt finance pipelineRunning in parallel, gated on agreement with the trial balance
Unattended model in the finance pipelineOff by default, deliberately
Client described rather than named, except where cleared.