Your AI Answered in 12 Seconds. Did the Customer’s Problem Get Solved? Build the Promise-to-Outcome Scorecard.
Your AI Answered in 12 Seconds. Did the Customer’s Problem Get Solved? Build the Promise-to-Outcome Scorecard.

Fast replies are easy to count. Kept promises are harder. Here is a two-Oppy system for measuring the difference.
A resident asks when the plumber is coming. A borrower asks whether her pay stubs arrived. A buyer asks if Friday’s showing is confirmed. A patient asks whether the referral made it to the dental office.
The AI answers in 12 seconds. It updates the CRM. It creates a task. The dashboard glows green.
Two days later, the customer asks again.
The system completed an action. The business did not complete the promise.
That distinction is becoming urgent. In a LinkedIn post published inside the past 24 hours, Commercial Observer technology partnerships director Edward Cohen wrote:
“Every multifamily operator is getting pitched AI right now. Very few are getting straight answers about what’s actually working.” 1
The straight answer will not come from message volume, average response time, CRM touches, or tickets closed. Those measures describe motion. They do not prove resolution.
A better question is brutally specific:
Did the customer receive the exact outcome the business promised, within the promised window, with evidence, accountable ownership, and no avoidable reversal?
This article shows how to answer it with a Promise-to-Outcome Scorecard, a two-Oppy operating design, five clocks, two exact prompt specifications, and a 30-day pilot.
The new AI stack can act. Your measurement system must catch up.
The September 10 AI product cycle made the direction plain. OpenAI announced a Data agent for governed enterprise analysis and approved actions, an Agents API for long-running tool-using work, and a full-duplex voice model for customer conversations. AWS announced durable agent execution that can resume after interruption without repeating completed model and tool calls. Microsoft published an AI Agent ROI Framework that includes success, resolution, and escalation rates, not merely productivity. These are provider announcements, not independent proof of customer outcomes, but they show where the operating stack is headed. 2 3 4 5 6
The business case remains less mature than the product surface. BCG reports that only 28% of 180 customer-service leaders in its worldwide survey had unlocked measurable value from generative AI. It also reports that 98% of executives view change management as crucial and 50% call it the main barrier to full value. Those figures are cross-industry survey findings, not residential-service benchmarks. They are still a useful warning: installing an agent is not the same as redesigning the work. 7
Real estate has little room for comforting fiction. HousingWire reported that average property insurance reached a record $209 per month in the second quarter of 2026, equal to 9.6% of the average mortgage payment. A vague “we are checking on it” costs more when customers are already absorbing higher ownership costs. 8
Even the real estate conversation is widening. In a First American podcast episode published inside the past 24 hours, Chief Economist Mark Fleming asked, “How do new economic constraints change where people live, work, and invest?” 9 Service operations face the same discipline. New tools matter only when they change an observable customer outcome.
Define a promise before you measure one
A promise is a customer-facing commitment that can be stated plainly and tested against an approved outcome definition.
“I’ll take care of it” is not a usable promise.
“I will confirm whether the showing is available by 4 p.m.” is.
The first phrase has no test, deadline, or owner. The second can be matched to a calendar record, a notification event, a timestamp, and a person responsible for exceptions.
This matters across residential service businesses. A mortgage team may promise a document-status update, but not approval. A title office may promise to confirm receipt, but not clear title. A property manager may promise a maintenance update, but not make a habitability judgment. A dental office may confirm receipt of a referral, but not diagnose the patient. A legal office may route intake documents, but not let an AI opine on the case.
The promise should be narrow enough to prove and modest enough to keep.
The Promise Loop
The Promise Loop follows one customer commitment through five stages. It works whether the promise originated with an Oppy, a human, or a mixed workflow.
| Stage | The operating question | Required evidence | Failure signal |
|---|---|---|---|
| 1. Capture | What did the customer request, and what did the business explicitly commit to do? | Exact quotes, channel, timestamp, case ID, due window, approved promise class. | Vague language, invented deadline, or no testable outcome. |
| 2. Assign | Who owns the next accountable step, and did that person or queue accept it? | Named role or owner, acceptance event, service class, escalation route. | A notification exists, but nobody accepted the work. |
| 3. Verify | What source-system event proves the work occurred? | Calendar, ticket, document, work-order, call, or case event with timestamp and source ID. | The ticket closed without outcome evidence. |
| 4. Confirm | Was the customer informed, and did the outcome remain stable? | Delivery or receipt event where available, follow-up window, repeat-contact check. | Silence is treated as satisfaction. |
| 5. Learn | Was there a reopen, correction, complaint, opt-out, missed deadline, or override? | Reason code, policy and prompt versions, corrective owner, repair timestamp. | Fast closures hide repeated work. |
The loop does not claim to measure universal customer satisfaction. It measures whether a defined operational promise was supported by the required evidence. That is narrower. It is also far more defensible.
Run five clocks, not one stopwatch
Average response time is useful. It is also easy to flatter.
A team can reply quickly, transfer slowly, work slowly, confirm poorly, and repair errors glacially. One stopwatch collapses those failures into a pleasant average. Use five.
| Clock | Start and stop | What it reveals |
|---|---|---|
| Promise clock | Customer request to a clear business commitment. | How quickly ambiguity becomes a testable next step. |
| Ownership clock | Promise made to accountable owner acceptance. | Whether handoffs create responsibility or merely notifications. |
| Resolution clock | Owner acceptance to evidence-backed completion. | Where the actual work stalls or repeats. |
| Confirmation clock | Completion to customer confirmation or the end of a defined no-reopen window. | Whether closure exists from the customer’s side. |
| Correction clock | Reopen, complaint, opt-out, or error detection to documented repair. | How fast the organization learns when it is wrong. |
Then calculate one hard headline metric:
Promise-kept rate = eligible promises completed with required evidence, on time, and without a defined reversal during the observation window, divided by all eligible promises due in the period.
Do not quietly delete difficult work from the denominator. A cancelled case should leave a recorded cancellation reason. Missing evidence should remain missing. Unknown is not passed. Dashboards enjoy optimism. Customers usually prefer plumbing.
Build two Oppies with incompatible powers
Oppy publicly describes custom AI employees that can work across phone, email, SMS, and integrated business tools. Its public functions guide recommends least privilege and documents a needs_attention workflow that can hold outgoing messages while a human responds. Its Account Concierge documentation also describes five event reports with a default comparison of the latest 30 days against the preceding 30 days. 10 11 12
Those capabilities make a 30-day outcome pilot practical, but the Promise Registrar and Outcome Auditor below are an operating design, not claims of prebuilt Oppy templates or native reports.
| Role | May do | Must not do |
|---|---|---|
| Promise Registrar Oppy | Read approved conversations and records. Extract explicit requests and commitments. Draft a structured Promise Record. Flag missing fields. Route ambiguity to needs_attention. |
Invent deadlines, decide eligibility, interpret coverage, make credit or tenant decisions, provide legal or clinical judgment, or send an unapproved commitment. |
| Outcome Auditor Oppy | Read approved Promise Records and source events. Test evidence, timing, reopens, and corrections. Produce passed, failed, pending, or human-review status with source links. | Modify source records, close consequential cases, suppress complaints, judge legal compliance, or change a score to improve performance. |
| Human Outcome Owner | Accept work, judge sensitive evidence, perform consequential actions, communicate exceptions, and approve changes. | Delegate licensed, legally reserved, clinical, underwriting, fair-housing-sensitive, or discretionary decisions to the scorecard. |
The separation is deliberate. The agent that drafts the promise should not certify its own success. Otherwise an automation’s note can become evidence that the same automation was brilliant. That is performance theater with JSON.
Prompt 1: Promise Registrar Oppy
Use this as a starting specification. Replace every bracketed item with an approved source, policy, and field. Do not treat it as legal advice or a ready-made production template.
ROLE
You are Promise Registrar. You convert approved customer interactions into proposed Promise Records.
APPROVED SOURCES
1. [Conversation transcript and message metadata]
2. [Named CRM record]
3. [Named scheduling, ticket, work-order, document, or case record]
4. [Approved service catalog and promise-language policy version]
Do not use memory, internet content, inferred personal traits, or any unlisted source.
TASK
Identify only an explicit customer request and an explicit business commitment supported by the approved sources. Produce one record per distinct commitment.
REQUIRED FIELDS
customer_request_quote
business_promise_quote
promise_category
case_id
customer_id
made_at_utc
due_at_utc, or null if no deadline exists
outcome_definition from the approved catalog
required_evidence
accountable_role
evidence_pointers
confidence: high, medium, or low
needs_human_review: true or false
HARD RULES
Never invent a deadline, owner, consent status, completion event, qualification, coverage interpretation, legal conclusion, clinical conclusion, credit decision, or tenant decision.
Never call a task complete because an outbound message was sent.
Route every missing, contradictory, sensitive, or out-of-catalog record for human review.
Preserve the customer’s exact complaint or revocation language. Do not judge legality.
Return JSON only.
Prompt 2: Outcome Auditor Oppy
ROLE
You are Outcome Auditor. You evaluate one approved Promise Record against read-only source-system events.
RETURN EXACTLY ONE STATE
passed: all required evidence exists, timing is met, and no defined reversal occurred.
failed: required evidence is absent, the due window was missed, or a defined reversal occurred.
pending: the due or observation window has not ended.
human_review: records conflict or the outcome requires qualified judgment.
HARD RULES
Treat an AI note, CRM field, or ticket closure as insufficient unless the approved outcome definition explicitly accepts it as evidence.
Never modify a source record, customer communication, score, compensation decision, staffing decision, or policy.
Cite every conclusion with source event IDs, timestamps, and missing-evidence reasons.
Never infer customer satisfaction from silence.
Escalate complaints, opt-outs, protected-class concerns, legal issues, clinical issues, coverage questions, underwriting matters, tenant decisions, and suspected misrouting.
Return JSON only.
Choose one promise class by business
| Business | Safe pilot promise | Evidence-backed outcome | Boundary that stays human |
|---|---|---|---|
| Brokerage | “I will confirm whether the showing is available.” | Calendar or showing system confirms status and the customer receives the update. | Advice, pricing, negotiation, agency, and fair-housing-sensitive guidance. |
| Mortgage | “I will identify which requested documents are still missing by 4 p.m.” | Approved checklist and receipt events support a sourced status message. | Approval, underwriting, credit, rate-lock, and legal interpretation. |
| Insurance | “I will route your declaration-page question today.” | The authorized queue accepts the case and contact is documented. | Coverage advice, comparison, binding, and eligibility. |
| Auto service | “I will confirm whether the replacement part arrived.” | Inventory or service-order event verifies status and the customer is notified. | Safety judgment, repair authorization, and warranty interpretation. |
| Title or escrow | “I will confirm receipt of the requested document.” | Receipt record and required notification both exist. | Title clearance, escrow decision, settlement approval, and legal interpretation. |
| Transaction coordination | “I will send the deadline checklist and flag missing signatures.” | Versioned checklist and signature-status evidence exist. | Contract interpretation and legal advice. |
| Property management | “I will update you on the maintenance request by tomorrow.” | Work-order status is verified, an owner accepts it, and the resident receives an approved update. | Habitability, emergency, lease, tenant, and enforcement decisions. |
| Dental office | “I will confirm receipt of the referral and offer the next available appointment.” | Referral and scheduling records support the response. | Diagnosis, treatment, urgency, and benefit interpretation. |
| Legal office | “I will confirm receipt of your intake documents and route them for review.” | Intake receipt and assigned review queue prove the step. | Legal advice, conflicts, merits, and deadline calculations unless counsel validates them. |
The scorecard
Start with six measures. Segment them by promise class, channel, risk level, team, and source system. Do not impose one universal benchmark across unlike work.
| Measure | Test | Management question |
|---|---|---|
| Promise-kept rate | Evidence-complete, on-time, no defined reversal divided by eligible promises due. | Did the business do what it said? |
| Unowned-promise rate | Promises without accepted ownership by the threshold divided by all promises. | Where are notifications masquerading as handoffs? |
| Evidence-complete rate | Records containing every required source pointer divided by records reviewed. | Can completion be proved? |
| Reopen rate | Outcomes returning with the same issue during the observation window divided by closed outcomes. | Are fast closures creating repeat work? |
| Correction clock | Median time from error, complaint, or opt-out to documented correction. | How quickly does the system repair failure? |
| Human override pattern | Count and reason codes for human changes to AI classification, routing, or status. | Which data, policy, or prompt assumptions are wrong? |
Microsoft’s new ROI framework likewise recommends measuring success, resolution, and escalation, along with governance and operating cost. The Promise-to-Outcome Scorecard narrows that logic to one service commitment and follows it until the evidence stops. 6
A 30-day launch plan
Do not begin with the whole customer journey. Begin with one boring promise. Boring is measurable. Measurable is improvable.
| Days | Operator action | Oppy contribution | Human gate |
|---|---|---|---|
| 1 to 5 | Pick one low-discretion promise. Define the outcome, source of truth, due window, cancellation rules, and reversal window. | Inventory recurring language from approved examples. | Legal, compliance, licensed, or clinical owner excludes prohibited promises. |
| 6 to 10 | Map conversation, CRM, calendar, ticket, document, and case events. Name the owner and escalation route. | Registrar drafts records from a small historical sample. | Manager validates every outcome definition and evidence source. |
| 11 to 15 | Run in shadow mode. Do not change customer communications or close cases from AI output. | Registrar and Auditor produce draft records and exceptions. | Two reviewers compare judgments and log disagreements. |
| 16 to 23 | Permit only approved tags or routing. Keep ambiguous and consequential cases in needs_attention. |
Registrar creates approved drafts. Auditor remains read-only. | Named owner accepts each exception and approves customer-facing action. |
| 24 to 30 | Review 20 to 50 records. Publish the five clocks. Inspect every failure and reversal. | Prepare the scorecard and evidence sample. | Cross-functional owner approves prompt, policy, staffing, or integration changes. |
BCG recommends evaluation-driven development and production monitoring for agentic workflows, including groundedness, routing accuracy, exceptions, latency, and drift. Its regulated-industry guidance also calls for explicit ownership, least privilege, traceability, and human intervention. Those are cross-industry principles, not an Oppy case study, but they fit this pilot cleanly. 13 14
Keep communications compliance separate from outcome measurement
A satisfied-looking scorecard is not a legal defense.
The Federal Communications Commission says current AI technologies that generate human voices fall within the Telephone Consumer Protection Act’s restrictions on artificial or prerecorded voice. For covered communications, reasonable revocations of consent must be honored within no more than 10 business days. The federal one-to-one seller consent rule was vacated and is not in force, but that does not make vague consent prudent. 15 16 17
A2P 10DLC registration is different. It is a carrier operating regime for application-to-person messaging over U.S. ten-digit long codes. Brand registration, campaign registration, STOP and HELP behavior, and campaign alignment affect delivery. They do not establish legal consent or prove TCPA compliance. 18
For the Promise Loop, treat a complaint, opt-out, or repeated contact as a signal requiring human review and correction. Do not let the Auditor declare a legal violation. Do not let the Registrar expand a campaign purpose, sender identity, or consent scope because a prompt sounded persuasive.
The managerial question worth keeping
An AI employee can answer, classify, summarize, schedule, route, and update records. Those are useful acts. None deserves to be called resolution until the promised customer state exists and the business can show why it believes so.
The next serious AI dashboard will not begin with “messages sent.” It will begin with:
- What did we promise?
- Who accepted responsibility?
- What proves completion?
- Did the customer have to ask again?
- How quickly did we repair what failed?
Twelve seconds is impressive. Keeping the promise is the product.
The Promise Loop, Promise Registrar, Outcome Auditor, prompts, and scorecard in this article are operating designs, not claims of prebuilt Oppy templates or native reports. This article is not legal advice. Product availability, communications rules, and state requirements should be verified for the relevant account, workflow, and jurisdiction.