Your AI Agent Did Exactly What You Allowed. Build Its Loss Budget Before the Next Tool Call.
By Anna with Oppy
Insurance has a tidy vocabulary for a hacker with stolen credentials. It has a more interesting problem when an AI employee causes damage with credentials the company intentionally supplied.
On August 27, Reuters reported that cyber insurers are reviewing policy language as autonomous agents take on more work. Traditional cyber coverage commonly anticipates a security event such as unauthorized access. An agent can create a costly outcome without either ingredient.1
“Some losses caused by AI agents will absolutely fall within cyber policies. The harder cases are where there is no conventional attacker and potentially no unauthorized credential use.”
Karthik Ramakrishnan, CEO and founder of Armilla AI, speaking to Reuters on August 27, 20261
The sentence every operator should keep is no unauthorized credential use.
Your AI did not break in. You invited it, connected the calendar, opened the inbox, handed it a CRM key, and taught it to be helpful. The loss can still be real.
A permission answers, “May the agent do this?” It does not answer, “How much can go wrong before a human notices?”
That second question needs its own operating system.
Valid Credentials Can Produce an Invalid Outcome
A conventional access review asks which systems an employee or agent can enter. An agentic review must also ask what can accumulate after entry.
| Authorized capability | Useful action | Plausible authorized mistake | Exposure that can accumulate |
|---|---|---|---|
| CRM write access | Update contact status and notes | Overwrite a source field or close the wrong records | Hundreds of corrupted records before the next audit |
| Email access | Send requested updates | Send a correct attachment to the wrong approved contact | Privacy, contractual, and remediation costs |
| Calendar control | Book and reschedule appointments | Move an entire service day while resolving one conflict | Lost appointments, overtime, refunds, and churn |
| Transaction workspace access | Prepare documents and route signatures | Use an outdated form or attach it to the wrong file | Delayed closings and repeated corrective work |
| Maintenance dispatch | Assign approved vendors | Dispatch duplicate visits while retrying a failed confirmation | Trip charges, tenant frustration, and budget leakage |
| Policy or quote workflow access | Collect data and request indications | Reuse stale facts across multiple submissions | Rework, disclosure risk, and unreliable recommendations |
None requires a cinematic robot rebellion. A retry loop and a Friday afternoon will do.
Reuters separately summarized controlled testing incidents disclosed by major model developers in July and August. Configuration errors or unintended internet access allowed agents to reach real third-party systems. No damage was reported, but the tests showed why a sandbox, credential boundary, and external target list must be controls rather than decorative nouns.2
The question is not whether a model is “safe.” The useful question is narrower: What is the maximum loss this specific agent can create with its current tools, rules, credentials, and time before detection?
Real Estate Is Moving From Answers to Actions
The timing is not theoretical. Inman reported on August 27 that new real estate integrations can let AI assistants interact directly with brokerage transaction data through the Model Context Protocol, rather than merely answer questions about it. The described workflow extends toward opening a transaction, preparing paperwork, and routing it for signature through conversation.3
A mortgage podcast published the same day made the permission expansion wonderfully plain. During Lykken on Lending’s August 27 episode, Jennifer Vallinayagam described a personal agent this way:4
“It has access to my calendar. It has access to my inbox because it also helps me manage my email inbox.”
Jennifer Vallinayagam, Lykken on Lending, August 27, 20264
That is the bargain. Convenience grows by connecting tools. Exposure grows by connecting the same tools.
The scale is arriving too. An August 19 Partners Group announcement says Emeria is deploying agentic AI across a property-management platform covering three million dwellings, seven countries, and 700 agencies. The company projects an initial EBITDA margin uplift of about 130 basis points. That is a projection, not a realized result, but the operating direction is unmistakable: more workflows, more speed, and more actions per employee.5
BCG’s current COO playbook argues that scaled agentic operations require common data structures, security rules, architectures, and explicit decision rights. Narrow pilots can work alone and still fail when connected.6
An agent that can only draft a reply has a quality problem. An agent that can draft, send, update, book, dispatch, and retry has a balance-sheet problem too.
Build an Autonomy Exposure Schedule
Create one Autonomy Exposure Schedule for every production AI employee. It is an auditable inventory of what the agent can do, how loss can accumulate, when a human must intervene, and what evidence will remain afterward.
This is not the same as an action log.
An action log tells you what happened. An Action Receipt can prove the checks performed before one event. The Autonomy Exposure Schedule looks forward and asks how large the outcome can become across many individually permitted events.
| Schedule field | Question it must answer | Example |
|---|---|---|
| Agent and job | Which narrowly defined AI employee is acting? | Property Maintenance Dispatcher |
| System and credential | Where can it act, and through which identity? | Work-order system, service account oppy_dispatch_01 |
| Data touched | Which records can it read, write, export, or send? | Tenant contact, unit, issue, vendor availability |
| Irreversible action | What cannot be reliably undone? | Vendor arrival, message sent, document disclosed |
| Maximum single-action value | What is the largest approved commitment per action? | $250 diagnostic visit |
| Maximum accumulated value | What is the largest exposure per hour and per day? | $750 per hour, $2,000 per day |
| Detection window | How long can bad behavior continue before a signal fires? | Five minutes or three similar actions |
| Reversal window | How long can a human undo the action? | Ten minutes before dispatch release |
| Human gate | Which condition requires named approval? | After-hours entry, safety issue, or cost above $250 |
| Kill switch | Who can suspend the tool or credential immediately? | On-call property manager and operations lead |
| Evidence retained | What will an investigator, customer, or insurer receive? | Prompt version, sources, tool call, approval, response, timestamp |
| Coverage question | Which policy and clause need a written answer? | Cyber, technology E&O, crime, professional liability, or general liability |
| Owner and test date | Who owns the exposure, and when was it last tested? | VP Operations, tabletop completed 2026-08-21 |
The companion control is the Loss Budget. This is the approved ceiling for cumulative authorized exposure, not a target and certainly not permission to lose it.
Every Oppy should have three budgets:
| Budget | Purpose | Trigger |
|---|---|---|
| Action budget | Limits one commitment, disclosure, dispatch, update, or transfer | One proposed action exceeds its value or risk class |
| Velocity budget | Limits repeated actions in a short period | Count, value, recipient, or destination changes too quickly |
| Daily exposure budget | Limits cumulative downside before review | Total commitments or affected records reach the approved ceiling |
The crucial control is automatic deceleration. When an agent nears a budget, it should reduce batch size, require confirmation, switch to draft-only mode, or stop. Sending a louder alert while the workflow continues is merely an enthusiastic incident report.
Give Four Oppies Incompatible Powers
Do not ask the production agent to grade its own risk. Build four narrow Oppies around it.
| Oppy | Job | May do | Must never do |
|---|---|---|---|
| Capability Registrar | Maintain the live inventory of tools, credentials, actions, recipients, and data classes | Detect permission changes and demand an owner | Grant access or approve its own inventory |
| Loss Budgeter | Model single-action and cumulative exposure from approved business assumptions | Recommend caps, gates, and rollback windows | Change production permissions or interpret insurance coverage |
| Coverage Mapper | Convert the schedule into precise questions for counsel and a licensed insurance broker | Attach policy text, endorsements, definitions, and written responses | Promise that a loss is covered or excluded |
| Drift Warden | Compare actual tool use with the approved schedule | Slow, quarantine, or pause a workflow when thresholds break | Expand the budget, restore access, or close the incident |
Oppy’s public product pages describe custom AI employees operating across phone, email, SMS, and more than 70 integrated business tools, with roles spanning qualification, support, CRM work, scheduling, listings, and content.7 Oppy also documents custom handoff rules, calendar connection, meeting capture, searchable conversations, knowledge accumulation, and confirmation before proposed actions.8 9
Those capabilities make a practical control pattern possible: one Oppy performs the job, while other Oppies observe permissions, measure accumulated exposure, prepare human decisions, and stop drift.
The separation matters. A cashier does not set the store’s credit limit while ringing the sale.
Use This Prompt for the Drift Warden
You are the Drift Warden for [ORGANIZATION].
Your only job is to compare actual agent activity with the approved
Autonomy Exposure Schedule. You may read tool-call events, action values,
recipient and destination changes, approval records, error codes, retries,
and the current action, velocity, and daily exposure budgets.
You may not perform the customer-facing job. You may not grant access,
increase a limit, interpret insurance coverage, or restore a suspended
credential.
Before each monitored action, verify:
1. the agent identity and approved job,
2. the exact system, credential, action, and destination,
3. the single-action budget,
4. the rolling 15-minute velocity budget,
5. the daily accumulated exposure,
6. the required human gate,
7. the reversal window and kill-switch owner.
If any field is missing, changed, stale, or over budget, emit PAUSE.
If three similar actions fail or retry, emit PAUSE.
If the destination, recipient class, payment instruction, legal form,
licensed decision, or physical-access condition changes, emit HUMAN_REVIEW.
Return only:
status, reason_codes, observed_activity, approved_limit,
remaining_budget, evidence_links, owner, and next_allowed_action.
A small structured record keeps the control inspectable:
{
"agent_id": "oppy_property_dispatch_01",
"job": "maintenance_dispatch",
"tool": "work_order_system",
"action": "authorize_diagnostic_visit",
"single_action_value_usd": 225,
"rolling_15m_value_usd": 450,
"daily_value_usd": 1175,
"limits": {
"single_action_usd": 250,
"rolling_15m_usd": 750,
"daily_usd": 2000
},
"reversal_window_minutes": 10,
"required_human_gate": false,
"status": "ALLOW",
"evidence_links": ["event://evt_48391", "schedule://aes_v7"],
"owner": "property_manager_on_call"
}
The dollar fields are examples, not recommended thresholds. A brokerage, lender, title company, insurer, property manager, dental practice, law office, or repair shop should calculate limits from its own authority, contracts, customer commitments, licensing rules, systems, and coverage.
Calculate Exposure by Job, Not by Model
The model name is rarely the useful unit. The job is.
| Business | Agent job | Loss Budget should measure | Hard human gate |
|---|---|---|---|
| Mortgage | Document chaser | Recipients contacted, documents moved, deadlines represented | Eligibility, pricing, lock, adverse action, or licensed advice |
| Title | Closing coordinator | Files changed, disclosures sent, identity checks requested | Any payment destination or wire-instruction change |
| Insurance | Intake and renewal coordinator | Records reused, quote requests sent, follow-ups issued | Binding, coverage interpretation, cancellation, or licensed recommendation |
| Brokerage | Transaction assistant | Forms prepared, signatures routed, dates updated | Contract execution, legal interpretation, or material term change |
| Property management | Maintenance dispatcher | Vendor commitments, entries, resident messages, daily spend | Safety event, physical access, habitability issue, or cost above authority |
| Dental office | Scheduling and intake agent | Appointments moved, reminders sent, records requested | Diagnosis, clinical advice, medication, or treatment decision |
| Law office | Intake and deadline assistant | Matters opened, documents requested, dates calendared | Legal advice, representation decision, filing, settlement, or conflict waiver |
| Automotive service | Appointment and estimate coordinator | Slots booked, parts requested, estimates relayed | Safety diagnosis, repair authorization above limit, or warranty determination |
This design does not make regulated judgment autonomous. It makes the surrounding administration measurable, bounded, and easier for the right human to inspect.
Ask the Insurance Question Before the Incident
The Reuters report is not a declaration that a particular agent loss is covered or excluded. Policy wording, endorsements, exclusions, facts, jurisdiction, and notice all matter.1
Use the Autonomy Exposure Schedule to conduct a written review with qualified counsel and a licensed insurance professional. Ask precise questions:
| Question | Why it is better than “Are we covered for AI?” |
|---|---|
| How does each policy define a security failure, wrongful act, system failure, professional service, and authorized user? | Definitions decide which door a claim can enter |
| What happens if the agent uses a valid credential and no outsider gains access? | It isolates the Reuters scenario |
| Which coverage may respond to corrupted records, erroneous communications, business interruption, privacy events, or professional errors? | Different outcomes can implicate different policies |
| Are autonomous actions, model errors, vendor failures, or contractual liabilities excluded, sublimited, or subject to conditions? | “AI coverage” is rarely one clause |
| Does the organization have duties to test controls, preserve logs, obtain consent, notify promptly, or use named vendors? | Operational controls can affect claim handling |
| How are repeated low-value actions aggregated into one event, occurrence, claim, or retention? | Agent losses can accumulate quietly |
| What evidence should be preserved before, during, and after a pause? | The claims file should not begin with archaeology |
Get the response in writing and attach it to the schedule. “Our broker seemed comfortable on Zoom” has limited forensic charm.
A 14-Day Build
| Days | Work | Exit evidence |
|---|---|---|
| 1 to 2 | Inventory each Oppy, tool, credential, action, recipient class, and data class | Capability Registrar has one owner for every production permission |
| 3 to 4 | Identify irreversible actions and estimate single, velocity, and daily exposure | Finance and operations approve assumptions and units |
| 5 to 6 | Set human gates, deceleration rules, reversal windows, and kill switches | Every high-impact action has a named human and tested pause path |
| 7 to 8 | Build the Drift Warden in observe-only mode | It reproduces expected decisions on historical events |
| 9 to 10 | Run tabletop failures: wrong recipient, duplicate retry, stale form, mass reschedule, vendor overrun | The system pauses before the modeled Loss Budget is exhausted |
| 11 to 12 | Review policy language and written questions with qualified professionals | Coverage responses and open questions are attached to the schedule |
| 13 to 14 | Launch with the smallest budget and daily review | No unexplained permissions, orphaned credentials, or unowned exceptions remain |
Track six numbers from day one: maximum possible daily exposure, actual daily exposure, percent of actions reversible, median detection time, median time to suspend a credential, and number of permissions without a current owner.
If the first metric rises faster than useful work, the agent is not scaling. Its blast radius is.
The Best AI Employee Has a Spending Limit on Mistakes
Agentic AI is moving into the systems where residential businesses live: inboxes, calendars, CRMs, transaction files, quote workflows, work orders, and customer conversations. The economic case can be substantial. So can the unmeasured accumulation of ordinary errors.
Do not begin with the policy question, “Does our insurance cover AI?” Begin with the operational question an underwriter, broker, counsel, CFO, and COO can actually examine:
What can this Oppy do with valid credentials, how much can accumulate before detection, and which control stops it first?
Build the Autonomy Exposure Schedule. Set the Loss Budget. Give the Drift Warden no ambition beyond saying “pause” at exactly the right time.
Intelligence is useful. Bounded intelligence is employable.