P.
Pdro Brandão
AACI/v0.1/2026
Production AI Systems

AI AgentContinuous Improvement.

A practical framework for observing, evaluating, improving, and connecting production AI agents to business value.

§ 02 / 05
The Problem

Is it working?

Reliability / Operations

Is it good?

Quality / Evaluation

Is it creating value?

Outcomes / Economics

Evidence for one does not automatically answer the others.

§ 03 / 05
The System

From production signals
to controlled improvement.

Production Case / Intimação Pro

Where failure taxonomy, versioned changes and regression evaluation became part of the improvement loop.

Diagnose → Prioritize → Hypothesize → Intervene → Validate → Release → Production

The goal isn't to make the LLM deterministic. It's to make the improvement process more measurable, controlled and repeatable.

A signal is not a diagnosis.

§ 04 / 05
Control

Control the system
around the model.

A /JUDGE HEALTH

Can you trust the evaluator?

An evaluator can look reliable in aggregate while failing exactly where it matters.

90% overall ≠ 90% where it matters

B /SAFE IMPROVEMENT

Did the change actually improve the system?

Production failures become regression cases. Changes become hypotheses. Evidence determines the release.

53.3% → 86.7%

C /ATTRIBUTION

How far can the evidence support the claim?

Better model behavior does not automatically mean better operations or business outcomes.

Quality → Operations → Business → Financial

D /ECONOMICS & CONTROL

Does the controlled system still make economic sense?

The cost of evaluation, supervision and control belongs inside the business case for automation.

Control cost ∈ business case

§ 05 / 05
Evidence

Built from production.
Expanded through research.

Intimação Pro
3,188documents processed
99.22%workflow success
53.3% → 86.7%evaluation accuracy

The framework was subsequently applied in an enterprise production deployment, moving from a bounded system to higher-volume agent operations — and changing the problem from improving one model interaction to controlling an operating system around AI.

AACI v0.1 distinguishes what was observed, what was designed,
and what remains a hypothesis.

Explore the evidence on GitHub

Running AI agents
in production?

If your team is struggling with evaluation, regressions, production quality, control, or understanding whether AI is actually creating value — let's talk.