AI engineering standards
AI engineering that compounds with every release.
AI amplifies the engineering beneath it. So every system we build runs on one loop, and every lesson production teaches becomes the next test.
- In writing
- Every hard-to-reverse choice filed as a decision record
- On your cases
- Tests and evaluations gate every release
- In production
- Service levels, traces and a runbook for every alert
Engineering recordSystems our engineers built and ran in production.
- 40,500requests a minute at peak on a consumer backend
- 10,000concurrent players on a real-time game backend
- −90%support tickets after a production support assistant shipped
- −40%datastore cost after a live migration to a better-suited engine
The standard
Every change runs the loop. No exceptions.
Google's 2025 DORA report calls AI an amplifier: it speeds up strong engineering and exposes weak engineering. So agents, models, data platforms and applications are all held to the same loop, and each stage leaves its evidence in your repository.
- 01
Decide
Decisions are written before they are built.
- A decision record for every choice that is hard to reverse
- The options, the trade-offs and the signal that would reopen it
- A spec before code, for engineers and coding agents alike
docs/adr/ · specs/
- 02
Build
Small changes, each reviewed by a second engineer.
- Trunk-based work on short-lived branches
- Every change reviewed before it merges, AI-drafted code included
- Infrastructure, dashboards and prompts kept as versioned code
Your pull-request history
- 03
Prove
Nothing ships on the strength of a demo.
- Unit, integration and end-to-end tests required to merge
- Evaluations on your graded cases gate every prompt, model and tool change
- Secret, code and dependency scanning on every commit
tests/ · evals/ · your CI logs
- 04
Ship
Every release can be undone.
- Signed builds with an SBOM and build provenance
- Feature flags and a canary first, then wider
- The rollback written before the rollout, and migrations that reverse
Release notes · your pipeline
- 05
Run
Run to service levels agreed with you.
- An SLO for each service, with an error-budget policy in writing
- Every request and agent step traced to the OpenTelemetry conventions
- An owner and a runbook for every alert that can page a person
runbooks/ · your dashboards
- 06
Learn
Production writes the next test.
- A blameless review after every incident: the cause, the fix, the control added
- Every failure and every surprising request becomes a graded case
- Model versions pinned, and the evaluations rerun before any change
reviews/ · evals/cases/
AI in production
How AI stays right after the demo.
A model that passed yesterday can drift today: a provider updates it, your data shifts, customers ask what nobody tested. We engineer for all three, and prove it on every release.
- Graded cases from your own workBuilt with your experts from real requests and real failures, then scored by code where an answer can be checked exactly, by a model judge calibrated against people, and by people where it matters most.
- Two suites, two jobsA regression suite held at full marks protects what already works. A capability suite measures what the next release has to learn.
- Reliability counted, never assumedCritical tasks run many times over. Beside the average we report pass^k, the share of tasks that succeed on every attempt.
- Pinned, then provenModel versions are pinned. A new model, prompt or provider reaches production only after the full evaluation reruns and it matches or beats the one in service.
- Cost engineered, like latencyCost per task is modeled before launch and watched after it, with caching, routing and batching tuned against it.
Every step of every run is traced and costed in your observability. Live runs are scored each day, and a failure becomes a graded case.
Code written with AI
Agents draft. Engineers decide.
In Veracode's 2025 study of more than 100 models, AI-generated code failed security tests 45% of the time. Our engineers work with coding agents every day, and every line those agents draft clears the bar a person's code must clear.
- Spec before codeEvery agent works from a written spec and plan, checked into the repository beside the code it produces.
- An engineer owns every lineWhoever merges a change has read it, run it and answers for it. Agent-written code gets no discount at review.
- The machines check the machineTests, static analysis and dependency scanning run on every change, however it was written.
- Your code stays yoursCoding tools run on enterprise terms with zero data retention, approved by you and never trained on your code.
Measured
What we measure, and report every month.
The five DORA delivery metrics, the service levels you set, the quality of the AI and what it costs to run, read from your own systems and reported on one page.
Delivery
The five DORA metrics, read from your pipeline
- Deployment frequency
- 38a month
- Change lead time
- 19 hcommit to production
- Change fail rate
- 3.1%of deployments
- Failed deployment recovery
- 42 minmedian
- Deployment rework rate
- 2.4%unplanned fixes
Reliability
Against the service levels you set
- SLOs met
- 12 of 12this month
- Error budget left
- 64%rolling 30 days
- Alerts with a runbook
- 100%of paging alerts
AI quality
On your graded cases and live traffic
- Regression suite
- 413 / 413held at full marks
- pass^5, critical tasks
- 97%right on every attempt
- Grounded answers
- 99.1%of live answers checked
Cost
What it costs to run, per unit of work
- Cost per task
- −6%month on month
- Prompt cache hit rate
- 71%of input tokens
- Cloud spend against plan
- −4%under the plan
Technology radar
What we adopt, trial and avoid.
Opinions, dated and stated in public, each with its reason. Only what we have run in production reaches Adopt, and the same practice runs on AWS, Azure, Google Cloud or your own data center.
Edition 2026.3 · September 2026
Our defaults. A system starts on these unless a constraint rules one out.
- Evaluation gates in CINo prompt, model or tool change merges without its evaluation run.
- Architecture decision recordsEvery hard-to-reverse choice keeps its reason on file.
- Spec-driven developmentCoding agents do their best work from a written spec and plan.
- Postgres with pgvectorRetrieval beside the records it cites, in a database your team already runs.
- Kubernetes on AWS, Azure or Google CloudPortable, well understood and easy to hire for.
- Terraform and OpenTofuEvery infrastructure change is a reviewed plan.
- GitHub Actions and GitLab CIThe pipeline lives in your repository, where you can read it.
- OpenTelemetryTraces, metrics and logs in any backend you choose.
- TypeScriptOne typed language from the browser to the API.
- PythonThe language of models, data and evaluation.
- Next.js and ReactFast, accessible interfaces a large pool of engineers can maintain.
Your system's decisions, on file.
Every choice that is hard to reverse gets a short record beside the code, so the next engineer, yours or ours, can read the reason.
- docs/adr/Every hard-to-reverse choice, with its reason
- specs/What each change must do, before it is built
- evals/Graded cases, judges and thresholds
- runbooks/One for every alert that can page a person
- reviews/Each incident: the cause, the fix, the control added
- dashboards/Service levels, traces and cost, as code
- AGENTS.mdHow coding agents work in this repository
# ADR 0012: Keep agent memory in Postgres
Status Accepted
Context The support agent needs memory across sessions:
past tickets, stated preferences, open actions.
Customer records already live in Postgres.
Options 1. A separate memory service
2. A vector database beside Postgres
3. Tables and pgvector in the primary cluster
Decision Option 3. Memory rows carry the customer key,
so deletion and retention follow the records.
Gains One system to secure, back up and audit.
Costs More load on the primary cluster.
Revisit Memory reads over 30% of primary load.Keep exploring
The engineering behind AI that runs the business.
Insights on engineering AI
- LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship
- MCP and A2A in the Enterprise: How AI Agents Reach Your Systems
- AI Inference Cost: How to Govern LLM and Agent Spend
Services built to this standard
Get in touch
Have your engineers talk to ours.
Write it as big as you imagine it.
15 answers, on the recordWhat CTOs ask before they hand over the systems.
The standard
What are AI engineering standards?
The practices that make an AI system dependable in production: decisions written down, every change reviewed, evaluations on real cases before each release, staged rollouts, service levels, tracing and incident reviews. Ours are on this page, and they apply to every system we build, AI or not.
Can you work to our engineering standards instead of yours?
Yes. Where your standards differ from ours, yours govern your systems, and the difference is recorded in a decision record so the next engineer knows why. Most clients keep their own repositories, pipelines and observability, and we work inside them.
Can our own team take the system over?
Yes, by design. The code, tests, evaluation sets, decision records, runbooks and dashboards live in your repository from the first commit, and the handover is written with acceptance criteria before the work starts, so your engineers inherit the reasons along with the system.
What is an architecture decision record?
A short document, kept beside the code, that records one decision: its context, the options considered, what was chosen, what it gains, what it costs and the signal that would reopen it. We write one for every choice that is hard to reverse, so no decision depends on anyone's memory.
AI quality
How do you evaluate an LLM or AI agent before release?
On graded cases drawn from your own work. A regression suite, held at full marks, protects what already works; a capability suite measures what the release has to learn. Answers are scored by code where they can be checked exactly, by a model judge calibrated against your experts, and by people where it matters most, and the suite runs in CI on every prompt, model and tool change.
What is pass^k, and why does it matter?
pass^k is the share of tasks an agent gets right on every one of k attempts. An average can hide an agent that is right four times in five; pass^k shows it. For critical tasks we report both, because a system a business runs on has to be right every time.
Can an LLM judge be trusted to grade AI output?
Only once it agrees with people. Each model judge is calibrated on cases your experts have graded, and rechecked when the judge or the task changes. Wherever an answer can be verified exactly, code grades it instead.
What happens when a provider updates the model?
Nothing reaches your users until it is proven. Model versions are pinned; a new version, prompt or provider runs the full evaluation first and ships only when it matches or beats the one in production, through the same staged release as any other change.
How do you monitor AI agents in production?
With traces of every run: each model call, retrieval and tool call, with its latency, tokens and cost, recorded to the OpenTelemetry GenAI conventions in the observability you already use. A share of live traffic is scored every day, and alerts fire when quality, latency or cost drift out of their band.
How do you keep the cost of running AI under control?
By engineering it like latency. Cost per task is modeled before launch and watched after it, and we tune it with prompt caching, routing routine steps to smaller models, batching and budgets that alert before a bill surprises anyone. The monthly scorecard reports it beside quality.
Delivery and reliability
What are DORA metrics, and do you report them?
They are five measures of software delivery from Google's DORA research program: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate. We read all five from your pipeline and report them every month.
What is an error budget?
The amount of unreliability a service level objective allows over a period. While budget remains, the team ships; once it is spent, reliability work goes ahead of new features until the service is back within its objective. The policy is agreed with you in writing, per service.
What happens when something breaks in production?
The alert reaches an engineer who owns it, with a runbook for that alert. We restore service first, keep you informed as it happens, and then write a blameless review: the cause, the fix and the control added. The failure also becomes a test, so the same fault cannot ship twice unnoticed.
Code written with AI
Is AI-generated code safe to ship?
Only after the checks any code must pass. In Veracode's 2025 study, AI-generated code failed security tests 45% of the time. So code a coding agent drafts gets the same review, tests, static analysis and dependency scanning as code a person writes, and an engineer who merges it answers for it.
Do your engineers use AI coding tools on our code?
Yes, with your approval, and only on enterprise terms with zero data retention, so your code is never used to train a model. Agents work from written specs, and every line they draft is reviewed, tested and owned by an engineer before it merges.
Not answered here? Two lines are enough.
Ask your own question