Working principles
How I work
Eight things I do the same way on every engagement, and the reasoning behind each one. They are what the case studies have in common once you strip out the domain.
Specification before generation
Prompting at a professional level means expressing intent precisely enough that a machine can execute it: goal, scope, constraints, edge cases, escalation conditions, logging, success criteria. Show me an original brief with acceptance criteria and I can tell whether you have the skill. Show me the finished product and I cannot. Every skill I ship starts as a spec, and the spec gets reviewed like code.
The system of record always wins
AI reads truth and drafts against it. It never recomputes or overwrites the source of truth. This is the most load-bearing choice I make. It keeps the AI layer out of the blast radius of the systems a business runs on, and it keeps me out of a turf fight with whoever owns those systems. Half of that sentence is architecture and half of it is organizational politics, and both halves are why the rule holds.
Human-in-the-loop as an architectural property
Every artifact carries a review line naming the human who signed off. Autonomy is a per-skill dial earned through telemetry, and each notch up is a cost the program pays for, never a feature it gets for free. When an executive asks for 99.5% draft acceptance and the comparable industry norm is 70 to 80%, I put the arithmetic on the table and we set the target where it lands.
The auditor writes a report and a human merges it
Governance runs as deterministic mechanical checks plus an LLM judgment pass that routes risk to named reviewers. Credit-adjacent work goes to Compliance. NPI goes to General Counsel. An external send with no human in the loop is a hard reject. Then a human merges the pull request, and the auditor never does. Take the human out of an irreversible action and what you have built is exposure with a governance label on it.
Evaluation is the job
AI is fluent and confidently wrong at the same time, which is why a polished output tells you nothing about whether it is right. Everything I ship carries the checks that caught its failures, and when a production skill breaks on a live desk the fix is a new rule in those checks. A cleverer prompt buys you about a week. I write pass/fail criteria somebody else can apply without me in the room. Observability is really two jobs on two backends: who uses a skill and at what cost is a separate question from why it produced a bad output.
Cost is a design input
Senior practice is deciding whether an AI workflow is worth running before scaling it. I derive cost per skill from real traces, even when it takes a session-id join across spans because under 1% of events carry both the cost and the name. The best model for a step is rarely the most powerful one.
Status language is load-bearing
Built means it runs in production under change control with a named owner. Drafted and prototyped are the two states before that, and I keep all three words apart in status reports, because the risk and audit functions read those reports and remember what the last one said. I would rather hand up the smaller number than defend an inflated one six weeks later.
What I am not
I do not train models. Frontier models plus retrieval get you about 90% of the value for a fraction of the cost, and the remaining 10% wants a real ML engineer, which I am not. I am not a prompt engineer either; that is the wrong altitude for the problems I take on. The work I do is platform architecture, governance, and getting things into production. If what you need is model training or prompt work, I will tell you on the first call.