# Agentic engineering at enterprise scale

Case study · Mars veterinary diagnostics

Agents that diagnose live production incidents, review pull requests, and draft the change-control paperwork behind a deploy, inside a Mars business running 50,000+ diagnostic orders a day. Production MTTR went from weeks to under 24 hours, and every other number here carries the methodology behind it.

Canonical: https://davidveksler.com/work/enterprise-agentic-engineering/
Updated: 2026-07
Primary artifact: https://cheatsheets.davidveksler.com/governing-agentic-ai.html?ref=portfolio

- 4.2x: sprint velocity (story points, same team and sizing, before vs. after)
- 5% → 80%+: code coverage in a sprint (CI threshold gate)
- 97% → 99.99%: platform uptime (quarter window), via the MTTR collapse
- 50K+: daily diagnostic orders
- weeks → <24h: production MTTR

---

## Problem

A veterinary diagnostics business inside Mars, running 50,000+ diagnostic orders a day, needed to modernize a large legacy platform and raise reliability without slowing feature delivery. The open question was whether agentic AI could carry real engineering load at that scale, under the same change control everything else answers to.

## My role

I lead agentic AI adoption in R&D. The direction of travel is code that is 100% agent-drafted, with the engineer's judgment spent on review and approval. Typing is the cheapest thing a senior engineer does and it is the part we gave away first. The 4.2x team was my direct team; I led it while driving adoption across R&D.

## What the agents actually do

- **Diagnose live production incidents** against Oracle and Application Insights, read-only, and hand a human the findings.
- **Run senior-level PR review** and draft the change-control paperwork that gates a deploy, each reading the system of record and handing a human the write approval.
- **Resolve production outages** and keep long-running multi-month tasks on track under supervision.
- **Generate tests in CI/CD**, which is how coverage went from 5% to over 80% in a single sprint.

Each of these keeps the same invariant as the [regulated-lender platform work](/work/regulated-lender-ai-platform/): the agent reads truth and drafts against it; a named human holds the write. The patterns and failure modes behind these workflows are written up in the [Agentic AI field guide](https://cheatsheets.davidveksler.com/agentic-ai.html?ref=portfolio).

## How the numbers were measured

Here is how each figure was measured, including where the measurement is weaker than the headline.

- **4.2x velocity** is sprint velocity in story points from the team's sprint tracking: the same team, the same sizing conventions, before and after the agentic workflows landed. Story points are a relative measure and I treat them that way. The claim is that this team completed 4.2x the sized work per sprint. It says nothing about your team.
- **Coverage from 5% to over 80% in a sprint** came from LLM-driven test generation wired into CI/CD, enforced by a coverage-threshold gate. The boundary matters: coverage tells you which lines execute under test and stays silent on how strong the assertions are. I have run no mutation-testing pass over the generated suite, so the coverage number and the CI gate are the whole claim.
- **Uptime from 97% to 99.99%**, measured over a window of one quarter or less, comes with a mechanism attached. The platform's availability problem was dominated by how long incidents lasted, and agents collapsed that. Read-only diagnosis against the production databases and telemetry turned multi-week investigations into same-day findings, and production MTTR fell from weeks to under 24 hours. Cut incident duration by that much at the same incident rate and the availability arithmetic follows.

## Where these numbers come from

Internal sprint and telemetry data at my current employer, which I cannot publish. None of them have been audited by anyone outside the company, and I would treat an outside claim like this one the same way.

## Limits

The velocity figure is a story-point measure, with the softness that implies. The generated test suite is gated on coverage; test quality itself has had no independent validation. And this describes work at a current employer, so I can give you the mechanism and keep the raw data private, which is the trade.

## Why it matters next to the regulated-lender case study

This is the same competence told from the opposite angle. The regulated-lender engagement shows AI as a governed product inside a regulated domain. This one shows AI carrying engineering load across a large org. One is founding the function; the other is running it at scale.
