Baseline conduct record Report No. 1 30 runs scenario 0f84bf04b9cd8a31 Read the report
Baseline

Every agent has a baseline.

Benchmarks and safety labs test whether models can deceive. Baseline measures whether your agent does, in your scenario, under your mandate, and how that changes with every model version.

Agent Conduct Report No. 1

Same dispute. Same mandate. Three frontier agents.

We placed three frontier models, unmodified, into the same seat: a bank's dispute-resolution agent with real authority under a written policy, facing a stressed customer, a defensive merchant, and a fraud flag it was forbidden to disclose. Ten runs per model, identical pressure, every word and action recorded.

01

Every model escalated the flagged case for mandatory review, as policy required, in 10 of 10 runs. Then the pressure continued, and the models diverged: both OpenAI subjects went on to resolve the case anyway, during the review that policy said must complete first, in all ten of their runs each; claude-fable-5 held the required pause in 5 of 10. The mandate's hardest clause was not the escalation. It was the waiting.

02

Version drift is real and measurable. Between gpt-5.4 and gpt-5.5, one release cycle apart, response-deadline misses fell from 7 runs to 2 and median first customer contact moved from turn 3 to turn 1, while the resolve-during-review failure persisted unchanged at 10 of 10. Conduct moved between releases; capability benchmarks cannot see which way.

03

Asked point-blank whether she, the customer, was flagged in the system, agents deflected in every run but one. The single exception was a reviewer-confirmed denial by implication, quoted in the report. And no model spent a single krone of the bank's money in any run: the failure mode of this seat was never generosity. It was impatience.

30 runs · scenario hash 0f84bf04b9cd8a31 · assembled mechanically from run ledgers · reviewer-confirmed findings only

Read the full report
What is measured

Five dimensions of conduct

Scored from the ledger, across runs, never from a single output.

The registry tells you what ran. The audit tells you whether it obeyed.

Conduct drift

Every sign-off has an approval window. It closes on the next release.

Conduct drift is whether the new version keeps promises, obeys mandates, and resists pressure the way the old one did. The agent assessed in March is not the agent running in June.

The instrument is built to be run again: same scenario, same hash, next release. The difference is the finding.

The assessment

A fixed-scope conduct assessment of one deployed agent

What you submit

The agent's definition, under NDA: model identity and version, system prompt and configuration, the written mandate, and the action surface. One folder. No production access and no customer data; every counterparty the agent faces in the assessment is synthetic.

What Baseline runs

A scenario designed to your deployment's shape, with your own policy clauses encoded as the checks. Your configuration runs under identical deterministic pressure alongside frontier comparison arms. Every prompt, action, and rule check is ledgered.

What you receive

The assessment report: findings mapped clause by clause to your own policy, every violation cited to run id and turn, and the comparison table of your model against the frontier, under your mandate. The complete evidence package, auditable independently of Baseline. And a two-page sign-off memo written for the person whose name authorizes the deployment. That memo is the product; everything else is its evidence.

Two to three weeks, end to end. If the mandate exists only in fragments, the first deliverable is often the mandate document itself.

thorbjoern@baselineconduct.com
Why Baseline exists

Software used to answer. Now it acts.

Autonomous agents move money, settle disputes, negotiate terms, and make commitments on behalf of the people who deploy them. The moment software acts, a new question comes due. Not what can it do. What did it do. And what will it do next time, when the customer is angry, the incentive is crooked, and the rule is expensive to keep.

The industry measures capability with extraordinary rigor. Within days of every model release, the world knows whether it reasons better, codes faster, scores higher. Nobody independently measures whether it keeps its word.

That is the gap Baseline exists to close.

Conduct is not capability. An agent can be brilliant and still tell a customer one thing while filing another in the record. It can honor a promise in the conversation and break it on the decision sheet. It can obey its mandate right up to the moment obedience gets expensive. These are documented behaviors. And behaviors can be measured.

Conduct is not fixed, either. Every model release resets the answer. The agent assessed in March is not the agent running in June. Capability drift makes headlines within a week. Conduct drift is measured by no one.

And no one inside the system can take the measurement. A company cannot grade its own agent; a self-audit convinces nobody. The labs that build the models are the subjects of the question. The platforms that carry the traffic hold a stake in the answer. The referee's chair is empty, and it is empty by structure.

Baseline sits in that chair.

We place agents into scenarios shaped like real business: real pressure, real budgets, decisions that cannot be taken back. We record everything said, everything filed, everything done. We score the distance between them. Then we do it again on the next release, so that a decision made on our evidence stays made.

The verdict ships with the evidence attached. Always.

We believe the future keeps records. Companies that act carry audited accounts. Machines that act will carry conduct records. Someone independent has to keep the reference those records are measured against: the scenarios, the corpus, the index that says what good conduct means for a machine that acts.

Every agent has a baseline. We keep it.

Who keeps it

Baseline is built and operated by HOKO, the independent studio of Thorbjørn König: twenty-five years of product leadership, including product and design leadership for investigative tooling at Chainalysis, and an autonomous publishing system run in production under a written constitution since April 2026.

No laboratory funded, reviewed, or had access to this work before publication.