Your AI Agent Did Exactly What You Allowed. Build Its Loss Budget Before the Next Tool Call.

By Anna with Oppy

Insurance has a tidy vocabulary for a hacker with stolen credentials. It has a more interesting problem when an AI employee causes damage with credentials the company intentionally supplied.

On August 27, Reuters reported that cyber insurers are reviewing policy language as autonomous agents take on more work. Traditional cyber coverage commonly anticipates a security event such as unauthorized access. An agent can create a costly outcome without either ingredient.1

“Some losses caused by AI agents will absolutely fall within cyber policies. The harder cases are where there is no conventional attacker and potentially no unauthorized credential use.”

Karthik Ramakrishnan, CEO and founder of Armilla AI, speaking to Reuters on August 27, 20261

The sentence every operator should keep is no unauthorized credential use.

Your AI did not break in. You invited it, connected the calendar, opened the inbox, handed it a CRM key, and taught it to be helpful. The loss can still be real.

A permission answers, “May the agent do this?” It does not answer, “How much can go wrong before a human notices?”

That second question needs its own operating system.

Valid Credentials Can Produce an Invalid Outcome

A conventional access review asks which systems an employee or agent can enter. An agentic review must also ask what can accumulate after entry.

Authorized capability Useful action Plausible authorized mistake Exposure that can accumulate
CRM write access Update contact status and notes Overwrite a source field or close the wrong records Hundreds of corrupted records before the next audit
Email access Send requested updates Send a correct attachment to the wrong approved contact Privacy, contractual, and remediation costs
Calendar control Book and reschedule appointments Move an entire service day while resolving one conflict Lost appointments, overtime, refunds, and churn
Transaction workspace access Prepare documents and route signatures Use an outdated form or attach it to the wrong file Delayed closings and repeated corrective work
Maintenance dispatch Assign approved vendors Dispatch duplicate visits while retrying a failed confirmation Trip charges, tenant frustration, and budget leakage
Policy or quote workflow access Collect data and request indications Reuse stale facts across multiple submissions Rework, disclosure risk, and unreliable recommendations

None requires a cinematic robot rebellion. A retry loop and a Friday afternoon will do.

Reuters separately summarized controlled testing incidents disclosed by major model developers in July and August. Configuration errors or unintended internet access allowed agents to reach real third-party systems. No damage was reported, but the tests showed why a sandbox, credential boundary, and external target list must be controls rather than decorative nouns.2

The question is not whether a model is “safe.” The useful question is narrower: What is the maximum loss this specific agent can create with its current tools, rules, credentials, and time before detection?

Real Estate Is Moving From Answers to Actions

The timing is not theoretical. Inman reported on August 27 that new real estate integrations can let AI assistants interact directly with brokerage transaction data through the Model Context Protocol, rather than merely answer questions about it. The described workflow extends toward opening a transaction, preparing paperwork, and routing it for signature through conversation.3

A mortgage podcast published the same day made the permission expansion wonderfully plain. During Lykken on Lending’s August 27 episode, Jennifer Vallinayagam described a personal agent this way:4

“It has access to my calendar. It has access to my inbox because it also helps me manage my email inbox.”

Jennifer Vallinayagam, Lykken on Lending, August 27, 20264

That is the bargain. Convenience grows by connecting tools. Exposure grows by connecting the same tools.

The scale is arriving too. An August 19 Partners Group announcement says Emeria is deploying agentic AI across a property-management platform covering three million dwellings, seven countries, and 700 agencies. The company projects an initial EBITDA margin uplift of about 130 basis points. That is a projection, not a realized result, but the operating direction is unmistakable: more workflows, more speed, and more actions per employee.5

BCG’s current COO playbook argues that scaled agentic operations require common data structures, security rules, architectures, and explicit decision rights. Narrow pilots can work alone and still fail when connected.6

An agent that can only draft a reply has a quality problem. An agent that can draft, send, update, book, dispatch, and retry has a balance-sheet problem too.

Build an Autonomy Exposure Schedule

Create one Autonomy Exposure Schedule for every production AI employee. It is an auditable inventory of what the agent can do, how loss can accumulate, when a human must intervene, and what evidence will remain afterward.

This is not the same as an action log.

An action log tells you what happened. An Action Receipt can prove the checks performed before one event. The Autonomy Exposure Schedule looks forward and asks how large the outcome can become across many individually permitted events.

Schedule field Question it must answer Example
Agent and job Which narrowly defined AI employee is acting? Property Maintenance Dispatcher
System and credential Where can it act, and through which identity? Work-order system, service account oppy_dispatch_01
Data touched Which records can it read, write, export, or send? Tenant contact, unit, issue, vendor availability
Irreversible action What cannot be reliably undone? Vendor arrival, message sent, document disclosed
Maximum single-action value What is the largest approved commitment per action? $250 diagnostic visit
Maximum accumulated value What is the largest exposure per hour and per day? $750 per hour, $2,000 per day
Detection window How long can bad behavior continue before a signal fires? Five minutes or three similar actions
Reversal window How long can a human undo the action? Ten minutes before dispatch release
Human gate Which condition requires named approval? After-hours entry, safety issue, or cost above $250
Kill switch Who can suspend the tool or credential immediately? On-call property manager and operations lead
Evidence retained What will an investigator, customer, or insurer receive? Prompt version, sources, tool call, approval, response, timestamp
Coverage question Which policy and clause need a written answer? Cyber, technology E&O, crime, professional liability, or general liability
Owner and test date Who owns the exposure, and when was it last tested? VP Operations, tabletop completed 2026-08-21

The companion control is the Loss Budget. This is the approved ceiling for cumulative authorized exposure, not a target and certainly not permission to lose it.

Every Oppy should have three budgets:

Budget Purpose Trigger
Action budget Limits one commitment, disclosure, dispatch, update, or transfer One proposed action exceeds its value or risk class
Velocity budget Limits repeated actions in a short period Count, value, recipient, or destination changes too quickly
Daily exposure budget Limits cumulative downside before review Total commitments or affected records reach the approved ceiling

The crucial control is automatic deceleration. When an agent nears a budget, it should reduce batch size, require confirmation, switch to draft-only mode, or stop. Sending a louder alert while the workflow continues is merely an enthusiastic incident report.

Give Four Oppies Incompatible Powers

Do not ask the production agent to grade its own risk. Build four narrow Oppies around it.

Oppy Job May do Must never do
Capability Registrar Maintain the live inventory of tools, credentials, actions, recipients, and data classes Detect permission changes and demand an owner Grant access or approve its own inventory
Loss Budgeter Model single-action and cumulative exposure from approved business assumptions Recommend caps, gates, and rollback windows Change production permissions or interpret insurance coverage
Coverage Mapper Convert the schedule into precise questions for counsel and a licensed insurance broker Attach policy text, endorsements, definitions, and written responses Promise that a loss is covered or excluded
Drift Warden Compare actual tool use with the approved schedule Slow, quarantine, or pause a workflow when thresholds break Expand the budget, restore access, or close the incident

Oppy’s public product pages describe custom AI employees operating across phone, email, SMS, and more than 70 integrated business tools, with roles spanning qualification, support, CRM work, scheduling, listings, and content.7 Oppy also documents custom handoff rules, calendar connection, meeting capture, searchable conversations, knowledge accumulation, and confirmation before proposed actions.8 9

Those capabilities make a practical control pattern possible: one Oppy performs the job, while other Oppies observe permissions, measure accumulated exposure, prepare human decisions, and stop drift.

The separation matters. A cashier does not set the store’s credit limit while ringing the sale.

Use This Prompt for the Drift Warden

You are the Drift Warden for [ORGANIZATION].

Your only job is to compare actual agent activity with the approved
Autonomy Exposure Schedule. You may read tool-call events, action values,
recipient and destination changes, approval records, error codes, retries,
and the current action, velocity, and daily exposure budgets.

You may not perform the customer-facing job. You may not grant access,
increase a limit, interpret insurance coverage, or restore a suspended
credential.

Before each monitored action, verify:
1. the agent identity and approved job,
2. the exact system, credential, action, and destination,
3. the single-action budget,
4. the rolling 15-minute velocity budget,
5. the daily accumulated exposure,
6. the required human gate,
7. the reversal window and kill-switch owner.

If any field is missing, changed, stale, or over budget, emit PAUSE.
If three similar actions fail or retry, emit PAUSE.
If the destination, recipient class, payment instruction, legal form,
licensed decision, or physical-access condition changes, emit HUMAN_REVIEW.

Return only:
status, reason_codes, observed_activity, approved_limit,
remaining_budget, evidence_links, owner, and next_allowed_action.

A small structured record keeps the control inspectable:

{
  "agent_id": "oppy_property_dispatch_01",
  "job": "maintenance_dispatch",
  "tool": "work_order_system",
  "action": "authorize_diagnostic_visit",
  "single_action_value_usd": 225,
  "rolling_15m_value_usd": 450,
  "daily_value_usd": 1175,
  "limits": {
    "single_action_usd": 250,
    "rolling_15m_usd": 750,
    "daily_usd": 2000
  },
  "reversal_window_minutes": 10,
  "required_human_gate": false,
  "status": "ALLOW",
  "evidence_links": ["event://evt_48391", "schedule://aes_v7"],
  "owner": "property_manager_on_call"
}

The dollar fields are examples, not recommended thresholds. A brokerage, lender, title company, insurer, property manager, dental practice, law office, or repair shop should calculate limits from its own authority, contracts, customer commitments, licensing rules, systems, and coverage.

Calculate Exposure by Job, Not by Model

The model name is rarely the useful unit. The job is.

Business Agent job Loss Budget should measure Hard human gate
Mortgage Document chaser Recipients contacted, documents moved, deadlines represented Eligibility, pricing, lock, adverse action, or licensed advice
Title Closing coordinator Files changed, disclosures sent, identity checks requested Any payment destination or wire-instruction change
Insurance Intake and renewal coordinator Records reused, quote requests sent, follow-ups issued Binding, coverage interpretation, cancellation, or licensed recommendation
Brokerage Transaction assistant Forms prepared, signatures routed, dates updated Contract execution, legal interpretation, or material term change
Property management Maintenance dispatcher Vendor commitments, entries, resident messages, daily spend Safety event, physical access, habitability issue, or cost above authority
Dental office Scheduling and intake agent Appointments moved, reminders sent, records requested Diagnosis, clinical advice, medication, or treatment decision
Law office Intake and deadline assistant Matters opened, documents requested, dates calendared Legal advice, representation decision, filing, settlement, or conflict waiver
Automotive service Appointment and estimate coordinator Slots booked, parts requested, estimates relayed Safety diagnosis, repair authorization above limit, or warranty determination

This design does not make regulated judgment autonomous. It makes the surrounding administration measurable, bounded, and easier for the right human to inspect.

Ask the Insurance Question Before the Incident

The Reuters report is not a declaration that a particular agent loss is covered or excluded. Policy wording, endorsements, exclusions, facts, jurisdiction, and notice all matter.1

Use the Autonomy Exposure Schedule to conduct a written review with qualified counsel and a licensed insurance professional. Ask precise questions:

Question Why it is better than “Are we covered for AI?”
How does each policy define a security failure, wrongful act, system failure, professional service, and authorized user? Definitions decide which door a claim can enter
What happens if the agent uses a valid credential and no outsider gains access? It isolates the Reuters scenario
Which coverage may respond to corrupted records, erroneous communications, business interruption, privacy events, or professional errors? Different outcomes can implicate different policies
Are autonomous actions, model errors, vendor failures, or contractual liabilities excluded, sublimited, or subject to conditions? “AI coverage” is rarely one clause
Does the organization have duties to test controls, preserve logs, obtain consent, notify promptly, or use named vendors? Operational controls can affect claim handling
How are repeated low-value actions aggregated into one event, occurrence, claim, or retention? Agent losses can accumulate quietly
What evidence should be preserved before, during, and after a pause? The claims file should not begin with archaeology

Get the response in writing and attach it to the schedule. “Our broker seemed comfortable on Zoom” has limited forensic charm.

A 14-Day Build

Days Work Exit evidence
1 to 2 Inventory each Oppy, tool, credential, action, recipient class, and data class Capability Registrar has one owner for every production permission
3 to 4 Identify irreversible actions and estimate single, velocity, and daily exposure Finance and operations approve assumptions and units
5 to 6 Set human gates, deceleration rules, reversal windows, and kill switches Every high-impact action has a named human and tested pause path
7 to 8 Build the Drift Warden in observe-only mode It reproduces expected decisions on historical events
9 to 10 Run tabletop failures: wrong recipient, duplicate retry, stale form, mass reschedule, vendor overrun The system pauses before the modeled Loss Budget is exhausted
11 to 12 Review policy language and written questions with qualified professionals Coverage responses and open questions are attached to the schedule
13 to 14 Launch with the smallest budget and daily review No unexplained permissions, orphaned credentials, or unowned exceptions remain

Track six numbers from day one: maximum possible daily exposure, actual daily exposure, percent of actions reversible, median detection time, median time to suspend a credential, and number of permissions without a current owner.

If the first metric rises faster than useful work, the agent is not scaling. Its blast radius is.

The Best AI Employee Has a Spending Limit on Mistakes

Agentic AI is moving into the systems where residential businesses live: inboxes, calendars, CRMs, transaction files, quote workflows, work orders, and customer conversations. The economic case can be substantial. So can the unmeasured accumulation of ordinary errors.

Do not begin with the policy question, “Does our insurance cover AI?” Begin with the operational question an underwriter, broker, counsel, CFO, and COO can actually examine:

What can this Oppy do with valid credentials, how much can accumulate before detection, and which control stops it first?

Build the Autonomy Exposure Schedule. Set the Loss Budget. Give the Drift Warden no ambition beyond saying “pause” at exactly the right time.

Intelligence is useful. Bounded intelligence is employable.

References