David Veksler
On this page

Case study · CRE bridge lending

Founding the AI function at a regulated lender

Two services were live in production at handoff, including a daily investor-enrichment pipeline, on a governed platform where nine departments could publish reviewed, drafts-only skills (agents propose, humans send). Two months, one engineer, on contract.

~345Ktrace events / month (measured)
3,685sessions in 30 days (measured)
$2.44model cost / session (measured)
30–80hrs/wk returned (measured at the workflow level, not a multi-week production curve)
9departments on one governed platform
2services live in production at handoff

Problem

A ~37-person regulated commercial-real-estate bridge-lending fund, nine departments, wanted AI to move deal and loan throughput. It had no AI function, no governance model, and a proprietary origination engine the CTO owned and protected. Getting a model to draft an email is a Tuesday. The thing that could actually hurt them was a credit-adjacent or NPI-leaking output reaching a borrower with no human in the path, in a domain governed by ECOA, FCRA, and SR 11-7.

Situation on arrival

There was an executive-authored five-phase AI roadmap and real top-down enthusiasm, with no delivery architecture underneath it. Production adoption was effectively zero: a chat tool used ad hoc by individuals, no shared skills, no governance, no observability, and no defined relationship between AI workflows and the systems of record. The roadmap assumed a march from literacy to autonomous agents to fine-tuned self-hosted models. It was missing the three things that decide whether such a program survives contact with a regulated lender: integration with the existing stack, a data-security and permission model, and fair-lending and model-risk controls.

The concrete problems followed from that. Deal truth, document truth, and file storage lived in three systems with no shared join key. There was no safe path for a non-engineer to ship automation without leaking credentials or fabricating numbers, and no way to see what any AI workflow had actually done, which is disqualifying for anything credit-adjacent. And the executive bar for autonomy, a 99.5% draft-acceptance threshold before removing human review, was being discussed as a near-term target rather than the long-horizon aspiration it is, against an industry norm closer to 70 to 80 percent for comparable sales-AI workflows.

Constraints

Two months, contract, one engineer. Reporting to the CTO with the CEO as executive sponsor. The system of record was off-limits to write. Every credit-adjacent artifact carried real regulatory exposure. And the metric I would be judged on was deal and loan throughput, which no model benchmark predicts.

My role

Principal AI Engineer, founding the function. I owned the technical architecture, skill and agent engineering, and integration with the existing lending stack. I did not own, and deliberately did not touch, the origination engine. My charter, confirmed directly with the CEO, was explicitly not AI literacy.

One constraint shaped everything else. This is a small company with no large engineering organization: the CTO owns the proprietary loan-origination engine that is the system of record, and that team plus me was essentially the entire technical bench. Every architectural decision was bounded by a single rule, that AI tooling builds on top of the core platform and never competes with it or forks its data.

Specification

“Done” meant a platform other people could safely build on: a place to publish a skill, an automatic audit that routed risk to the right human, a way to see who used what and at what cost, and a hard architectural guarantee that no model output reached a borrower without a named human signing off.

Vocabulary, briefly

Three words carry this case study, so here is what they mean outside the AI-tooling bubble. A skill is a versioned, reviewable unit of AI work: a specification plus the tools it may call, packaged so it can be published, audited, and measured like code. Plugins bundle skills for distribution. Connectors are the credentialed integrations a skill calls (the CRM, email, storage); the credential is the real access boundary.

Approach: platform before breadth

I built the rails before the traffic. A Git-backed skill marketplace distributed as an auto-synced plugin catalog. A publish-and-audit governance pipeline. An observability backend. Only then did department skills proliferate, and because the rails existed every one of them arrived versioned, audited, and instrumented on the way in. By handoff, ~30 skills had shipped across six of ten namespaces, all on the same rails.

The central decision was to treat the program as a platform problem rather than a collection of prompts. The unit of delivery is a skill: a versioned, reviewed instruction module with a declared safety classification and a fixed set of connector permissions. Skills are authored in a Git repository, distributed to the agentic client as an auto-synced plugin marketplace, and loaded on demand by description match. That gave me one place to enforce review, one audit trail, and a clean separation between the authors and the consumers of a capability. The alternative, letting each department accumulate private prompts inside the tool, is how a lender ends up with unreviewable shadow automation. I rejected it explicitly.

The spine runs authoring, then distribution, then execution, then observability. Governance sits at the authoring end as the publish-and-audit gate. The marketplace is the plugin catalog in the middle. The agentic client runs reviewed skills and drafts only. Traces flow out to the observability backend. The systems of record sit underneath the whole thing, read-only: the CRM, the origination engine, the document system, and file storage. Drafts go up to operators, never out to borrowers.

System architecture
authoring → distribution → execution → observability Governanceauthoring · gate Marketplaceplugin catalog Agentic clientruns reviewed skillsdrafts only ObservabilityOTel → backend Operatorshumans in the loop Interviewprototype AgentsHITL · queued Systems of record · read-onlyCRMOriginationDocumentsStorage drafts → human reads truth never overwrites merge loads traces queued
The platform, end to end. Solid is shipped, dashed is prototyped or specified. Every box links to the section that covers it.
shipped prototype specified

Key decisions

  • The system of record always wins. AI reads truth and drafts against it; it never recomputes or overwrites it. Every dashboard I built is a human-review aid, not a ledger. Accounts-payable aging, for instance, is assembled from invoice email for the weekly finance meeting, but the accounting system still owns the authoritative balance. This kept my layer out of the SoR’s blast radius, and out of competition with the CTO’s origination engine. The choice was as much about org boundaries as about architecture.
  • Human-in-the-loop lives in the architecture. Every artifact carries a review line naming the human who signed off. Outreach skills draft into the operator’s own mailbox; diligence skills produce a memo for a named reviewer. Autonomy is a per-skill dial earned through telemetry, so no single acceptance number gates the whole program. When the executive ask came in at 99.5% draft acceptance, I put the comparable sales-AI norm on the table, ~70 to 80%, and we set the target where the arithmetic put it.
  • The auditor writes a report and a human merges it. Deterministic mechanical checks (secrets and PII scan, frontmatter, connector-declared-equals-used, dedup) plus an LLM judgment pass that routes flags to named reviewers: credit-adjacent to Compliance, NPI to General Counsel, an external send with no human in the loop is a hard reject. Then a person merges the PR. The auditor never does.
  • Distribution and permission are separate systems. Installed skills are visible org-wide, so department segmentation is documentation and ergonomics only. The access boundary is the connector credential, held server-side. I organized departments as separate plugins for discoverability and ownership, and never relied on plugin membership for security.
  • Thin artifacts, fat services. Live dashboards can only make short calls to backing services, on the order of a sixty-second ceiling, so anything heavier runs as an asynchronous job behind a poll contract, with credentials held server-side in one place.
  • Observability is two jobs, not one. Who uses a skill and at what cost is a different question, on different infrastructure, from why a skill produced a bad output. I stood up a wide-event backend for the operational job because it ingests arbitrary telemetry with no schema gate, and reserved a purpose-built LLM-observability platform for the higher-stakes later-phase agents.

On the executive roadmap I pushed back on two points and carried both. Supervised fine-tuning on a self-hosted model should be conditional on earlier-phase performance data rather than assumed, because frontier models plus retrieval will plausibly cover ninety percent of the value at a fraction of the operational cost. And rather than let strategic priority override adoption readiness, I split the training tracks, sequenced by readiness, from agent investment, sequenced by strategic order, so neither blocked the other.

What shipped

Distinguishing what reached production from what I prototyped and what I specified in design. Department coaches authored many of the department-specific skills as forks of reference patterns I established; I owned the platform, the foundational skills, and the engineering review gate.

Platform and governance layer, in production. A Git repository deployed as a plugin marketplace: ten plugins enumerated through a manifest, one per department plus a shared plugin, with a CI pipeline that rebuilds the catalog and announces diffs to a team channel. A publish-and-audit governance path, where a publish step writes a candidate skill to a review folder and posts metadata to a helpdesk channel, which auto-files a triage ticket, and an auditor then runs mechanical checks and surfaces judgment flags with explicit reviewer routing. A chat-based helpdesk that converts member questions into tracked tickets. A standalone enrichment service running a research-and-drafting pipeline against contact cohorts, writing finished drafts back into the CRM behind a versioned JSON data contract, with a per-contact audit trail and a read-only health monitor. And a workflow-audit skill that produces a ranked list of automation candidates for an individual employee, with scope, estimated hours saved, and a recommended delivery pattern, which is the on-ramp by which non-technical staff find their first good target.

Department and shared skills, in production, drafts-only. Lending operations got a loan-document sorter that routes email attachments into the correct numbered deal subfolder, enforces recording numbers, and direct-messages the operator a gap list against the closing checklist, plus a prescreen credit-memo drafter, a credit-package-to-signature-envelope PDF converter, and a sponsor diligence and adverse-media research memo. SMB origination got a scheduled weekday job that drafts borrower follow-up emails for missing application fields, plus live pipeline-board and pre-quote dashboards over the CRM. Investor relations and growth got warm-outreach and dormant-broker re-engagement tooling, professional-network enrichment, and an accounts-payable aging dashboard for the weekly finance review. Productivity and brand got a per-user daily brief, recurring team priority boards, meeting action-item extraction, a session-handoff tool, and a brand-and-voice layer used as a final pass across other skills.

Prototype and design spec. A stakeholder interview bot, a branching-questionnaire web application over the model API producing structured plus narrative output, deployed for field research into automation priorities and not hardened for production. And a set of specified semi-autonomous agents: nightly portfolio surveillance, covenant monitoring, automated diligence assembly, and recurring investor-report drafting, each specified with human approval gates and a per-agent model card. An autonomous credit-memo agent sat queued behind resolving CRM write-scope gaps. Specified, not built.

One workflow on a real desk

The clearest test of the platform was taking one skill to a single operator and watching it survive contact. The skill was the scheduled weekday job that drafts borrower follow-ups for missing application fields. The operator was an SMB loan originator whose pre-quote book ran to roughly two dozen active deals, each needing the same unglamorous sweep: open the deal, work out which Stage 1 fields are still missing (entity name, EIN, term, estimated credit, and so on), and write the borrower an ask for exactly those. By hand that was an estimated five to ten minutes a deal, so a full sweep ran into hours and, in practice, slipped. Deals sat for days waiting on the same three or four fields.

Velocity mattered and the build showed it. The marketplace stood up, and within about two weeks four onboarding skills were tested and the audit-to-publish flow was dry-run validated. The first department-integrated workflow, this one, was in the SMB team’s hands for live testing roughly a week after that. First spec to team sign-off was sixteen days.

The first time it ran in front of the team, the skill was about to ask a broker for a borrower’s Social Security number. The model caught the ask and fell back to a neutral template on its own. The assumption behind it was still wrong, and that was the lesson.

It had assumed the primary CRM contact was the borrower. On the first pull, four of the first five contacts were brokers, and the draft logic was about to request personal items, an SSN and a date of birth, from a broker: the exact NPI-adjacent ask that torches a broker relationship. Separately, an internal audit stamp and a run of em dashes had leaked into borrower-facing copy, with drafts signing off “Assembled by the agent. Reviewed by [operator], [date],” which belongs in an audit log and not a borrower’s inbox. A reviewer’s QA pass the next day filed four more precise defects: contacts with a tax ID already on file were still being asked for an SSN, a deal with no borrower name rendered a broken subject line, and one contact the application API typed as a sponsor but the CRM’s lead-status field tagged as broker outreach exposed a source-of-truth conflict.

Each failure became a rule, and the integration got simpler rather than cleverer:

  • Dropped the direct lending-API call and its service token for the application tool, so there is no secret in the skill and no key to manage, and fixed a field-shape mismatch so the missing-field diff runs as written.
  • Made the deal-list query carry contact email and phone, so in degraded mode with the CRM offline the skill still drafts instead of skipping every deal for a missing recipient.
  • Skip broker-intermediated deals entirely, and treat the CRM’s lead-status field as authoritative over the application API’s type when the two disagree, so the skill never drafts a borrower ask to a broker.
  • Gate the SSN ask on the tax-ID-on-file flag so it stops asking for an SSN already on record, and add a borrower-name fallback so the subject line never breaks.
  • Strip the audit stamp and the em dashes from borrower-facing drafts, and move the stamp to the run summary where reviewers actually need it.

The measured delta: the first clean pass triaged the whole live pre-quote book in a single run of a few minutes, against a manual sweep measured in hours, and the originator’s sign-off was a one-line “let’s go.” The catch the telemetry forces me to state is that the skill did not then run for two sustained weeks of solo production. It was deliberately folded into a single interactive pre-quote artifact that became the one owner of morning drafting, to avoid duplicate drafts. So the measured reality is a validated workflow and a fast convergence onto the interactive surface, not a multi-week hours-saved curve.

Hard technical problems

Cross-system data join. Deal status lives in the CRM, document status in the lending document system, and the files in cloud storage referenced by path, with no shared primary key spanning them. I made the loan number embedded in the deal-folder title the canonical join key and built the document sorter to treat that token as the source of truth rather than trusting any single system’s record. That turned a three-way reconciliation problem into a deterministic lookup.

Connector permission model. The CRM connector reported write failures on specific object types through a distinct permission flag rather than a connection error. I diagnosed it as a per-object scope gap, with note and email writes blocked, and routed it to the platform owner as a narrowly scoped grant framed around agent activity-logging rather than blanket write access. The broader lesson, that the marketplace is not the access boundary and the connector credential is, hardened into a standing principle.

Enrichment integration without new infrastructure. Rather than build a bespoke adapter, I used CRM custom properties as the integration bus. A request flag, surfaced as a re-enrich control in the live dashboards, acts as a priority override into the enrichment queue, and results return as a single versioned JSON payload the consuming dashboard parses, with graceful fallback to a template. I trimmed the property schema from twelve fields to six, on the rule that a field exists only if the CRM itself must query, sort, filter, or display it, and everything else folds into the payload. The deterministic halves, candidate selection and write-back, are standard-library Python bracketing the model-driven core, and they compute the staleness hash from one shared module so the selector and the writer can never disagree.

Observability under a logging gap. The agentic client’s own activity is excluded from its audit logs, compliance API, and data exports, which is a real gap for anything credit-adjacent. I treated agent and skill execution as a telemetry problem and routed OpenTelemetry traces, spans, and outcomes into an independent backend so they are queryable regardless of the client’s native logging. That telemetry is the compensating control that makes credit-adjacent and scheduled workflows defensible. A secondary problem fell out of it: cost data lived on model spans while skill names lived on skill spans, and no span carried both. They share a session identifier, so a derived column joined on that identifier produced cost-per-skill-invocation with no new instrumentation.

Confused-deputy risk across tools. I threat-modeled whether a prompt injection delivered through untrusted content, an inbound email or a scraped page, could chain a draft-only email connector with browser automation to actually send a message. The mitigations were architectural rather than exhortation: isolate the sessions that process untrusted content from the sessions that can act on the mailbox so the chain cannot form, and place the human checkpoint at a layer the agent cannot bypass, the provider’s native send confirmation, rather than trusting the model to stop itself. Published platform evaluations put prompt-injection success at roughly 24% unprotected against roughly 1% with defenses on. Residual risk concentrates in well-camouflaged injections and in confirmation fatigue, which is why the checkpoint sits below the agent rather than inside it.

Sandbox-to-document-API authentication. An in-place document formatter needed a capability the storage connector could not perform, and the agent’s execution sandbox sits one trust layer below the connectors and never sees their OAuth tokens. I worked the four real options and recommended a service account as the immediate unblock, accepting the loss of per-user audit, with a custom protocol server as the clean answer to adopt before any further skills depend on that API. The discipline was separating unblock-testing-this-week from the-pattern-we-commit-to.

A second model for the jobs the primary agent could not reach. A couple of Google-native jobs sat outside the sandbox’s reach, and each got handed to the surface that already holds the data and the credentials. The agentic client can see an email’s attachments but cannot download the bytes, so a Gemini-powered Workspace Flow ingests every non-image attachment into a shared Drive folder, where the document sorter and the AP-aging dashboard pick them up. And because a full skill file or plugin bundle exceeds the Drive connector’s per-argument cap, the publish path writes through an Apps Script web app running as a service account, where the write credential already lives. The principle that generalizes: put each job next to the data it needs and let the primary agent orchestrate, rather than bending one platform to do everything.

Architecture principles established

  • The system of record always wins. AI reads truth and drafts against it; it never recomputes or overwrites it.
  • Drafts-only by default. A knowledgeable human validates every output against a known source of truth, and autonomy is earned per-skill through telemetry rather than granted up front. Every artifact carries a review line naming the human who signed off.
  • The marketplace is distribution, not permission. The connector credential is the real access boundary, and the design must reflect that.
  • Thin artifacts, fat services. Anything past the artifact timeout is an asynchronous job behind a poll contract, with keys held server-side in one place.
  • Observability is two jobs, not one. Operational analytics and agent-quality observability have different requirements, and raw loan-data traces stay out of any backend that does not redact at ingest.
  • Schema minimalism on shared systems. A field exists only if the system must query, sort, filter, or display it.
  • Precise status language is load-bearing. Drafted, prototyped, and built mean different things, because status inflation is how an AI program loses credibility with risk and audit.
  • Put the human checkpoint where injection cannot reach it. Design the boundary one layer below the agent rather than trusting the agent to stop itself.
  • Put each job next to its data and credentials. The primary agent orchestrates, but attachment ingestion runs as a Workspace Flow into Drive and oversized Drive writes go through an Apps Script service account.

How it was measured

I instrumented the whole thing with OpenTelemetry into a wide-event backend, ~345,000 events per month across the agent datasets, with cost-threshold alerting on top. That record also partially closed the platform’s compliance blind spot by creating a who-did-what-when history where none had existed.

Cost-per-skill was the hard part: cost lives on model spans, names live on skill spans, and under 1% of spans carry both, so a session-id join was the only path to a real number. That number: ~$8,980 of model spend over 30 days across 3,685 sessions, about $2.44 each. For scale, that $2.44 per session sits against a measured 30 to 80 hours per week returned across the firm. Human review time was the binding constraint the entire engagement, and model spend never came close, which is exactly what the per-skill autonomy dial exists to manage.

Adoption, by repeat-use cohort:

Repeat-use cohortShare of skills
Single use~26%
2 to 3 uses~32%
4 or more uses~42%

Recurring use concentrated in the authoring and platform-meta skills rather than the department skills. That is an early-rollout curve, not a finished one.

Hours returned per week, with the measured figure and the design target kept apart:

FigureRangeStatus
At observed adoption30 to 80 hrs/wkmeasured
Steady-state ceiling60 to 120 hrs/wkdesigned ceiling, not reached during the engagement

The range is gated on adoption and it is not a midpoint. The ceiling assumes full intended adoption, the telemetry confirms it was never reached, and I would not lead with it.

Tied back to the charter’s terminal metric of deal and loan throughput: on the one workflow taken to a real desk, a manual pre-quote sweep measured in hours collapsed to a single multi-minute run that drafted field-specific asks across the live book, and the credit-side prescreen and diligence skills target two to four analyst hours per deal. The throughput logic is that freed originator and analyst time converts into faster time-to-quote and more deals worked per head. I will not claim the closed-loan number, because the engagement ended before sustained production usage could show that conversion as a curve. The accurate statement is a validated time-to-value at the workflow level and an unproven, plausible link to throughput at the portfolio level.

Scope reached

The program reached an organization of roughly three dozen people. Ten plugins, a CI distribution pipeline, a chat-to-ticket governance and helpdesk path, a scheduled enrichment service with its own health monitoring, and on the order of thirty distinct skills, nineteen firm-authored. Those skills populate six of the ten plugin namespaces: CRE and lending operations, SMB origination, investor relations, finance, technology, and shared. The remaining four, marketing, legal, HR, and asset management, were scaffolded for ownership but not yet populated. The skills call seven first-class connectors, the CRM, email, cloud storage, calendar, team chat, project management, and a single managed lending API exposing both document upload and the application and admin reads, with the observability backend as an eighth connector that is a telemetry sink the skills never call. Zoom, Sheets, and browser automation appear in a handful of skills.

Governance and compliance

Because this is a regulated lender, I treated governance as load-bearing rather than paperwork. I argued for fair-lending controls, ECOA and FCRA considerations, alongside model-risk guidance, SR 11-7, in any credit-adjacent workflow, and made human-in-the-loop a property of the architecture so that no model output reaches a borrower or a credit decision without a named human in the path. The auditor flags credit-adjacent logic, financial calculations, writes to systems of record, and the handling of non-public personal information, routing each to the right reviewer rather than rubber-stamping it.

The auditor’s design point is that it produces a report, not a verdict. It runs in two halves: deterministic scripts handle the mechanical checks, so whether a regex matched a credential is never an LLM judgment call, and an LLM pass handles the judgment flags and routes each to a named reviewer with a blocking-or-advisory severity. The merge stays a human action against the pull request. If the model in the loop were the gate, the gate would be broken. It will auto-patch a narrow set of purely mechanical defects, frontmatter shape, connector-list coercion, naming, so schema nits do not cost a review round-trip, but it never edits workflow logic, never overrides a judgment flag, and never commits to the marketplace itself.

One data-protection caveat shaped the backend choice. The operational store ingests raw prompt text that can contain borrower names and deal addresses. That is acceptable for an internal analytics tool the AI team controls, and it is not the long-term home for agent traces once real loan data flows through later-phase workflows, where pre-redaction before storage is required.

The highest-leverage governance insight was structural. The brand-and-voice layer sits, by design, in the path of nearly every outgoing communication, which makes it the single best instrumentation point for a communications-layer compliance check: prohibited claims, disclosure language, ECOA-adjacent phrasing, and accidental disclosure of non-public pipeline detail, without building a separate compliance skill or relying on every author to invoke one.

What went wrong, and what I learned

Recurring use concentrated in the authoring and platform-meta skills: publish-skill, the skill-creator, the voice skill. The department value skills, which were the whole point, got used less. The hours-saved figure was validated at the workflow level over a short window, so nobody has watched it hold for a quarter. And the designed steady-state ceiling was a design target that two months never got near.

What worked was leading with platform and governance before breadth. Building the marketplace, the review gate, and the observability backend first meant that when department skills proliferated they landed on rails instead of as scattered prompts. Drafts-only as a default aged extremely well, because it defused the autonomy debate by turning autonomy into a per-skill dial. Using CRM custom properties as an integration bus shipped in days. The loan number as a canonical join key removed an entire category of reconciliation bugs. And quantifying everything moved leadership conversations from opinion to arithmetic.

What I would do differently: instrument adoption from day one rather than after skills shipped, so the hours-saved story rests on measured data. Force the connector permission model to an explicit decision earlier. Resolve system-of-record write scopes before designing any skill that depends on them. Push harder and sooner on separating the roadmap’s aspirations, autonomy targets and fine-tuning, from its near-term deliverables. And invest earlier in a catalog-level overlap reviewer rather than only the per-skill auditor, because overlapping authorship is the failure mode that quietly kills a skills platform and a per-skill check cannot see it.

Outcome

A governed AI platform a nine-department firm could build on, with measured cost and measured hours returned: 30 to 80 hours per week at observed adoption. At handoff, two services built on it were running in production, the chat-to-ticket helpdesk and a daily investor-enrichment pipeline. The pipeline gets its own case study, because it is the department-value work the adoption numbers above say was rare. The engagement ended when the fund took heavy redemptions and the planned conversion fell through. Nobody fell out over it.

Limits

The measured window is short. Two months is enough to found a function and prove the rails; a steady state takes longer than I had. The ceiling figure is a design target and it stays labeled as one everywhere it appears. I have no post-exit visibility into the fund’s systems, so everything here describes the state at handoff and nothing after it. The client and individuals are anonymized, and the internal telemetry behind the measured figures is not mine to publish, so the mechanism is here and the raw data is not.

Where it fits

The generalizable thesis: in a small, regulated, system-of-record-centric business, the highest-leverage AI work is not the cleverest agent. It is the platform that lets non-engineers ship reviewed, observable, drafts-only automation on top of the existing stack without ever competing with the system of record. Get the substrate right and capability compounds safely. Get it wrong and you accumulate unauditable risk that somebody eventually has to unwind.

This case study is founding the function. The Antech work is the same competence from the opposite angle, agentic engineering carrying load across a large organization rather than a governed product inside a regulated domain. The cheatsheets pipeline shows the same governance discipline at personal scale, where the spec, the merge gate, and the audit trail are all public and checkable.

Email me about this Where I fit best ← All case studies