Back to insights
16 September 20263 min readNexius Labs editorial

When AI refuses: lessons from Darryl Wong’s payment-workflow case study

A Nexius editorial summary of Darryl Wong’s When AI doesn’t comply: what changed, what remains unknown, and how teams can continue safely.

Original reporting by Darryl Wong. Read “When AI doesn't comply

Visual concept: When AI refuses: lessons from Darryl Wong’s payment-workflow case study within a human-controlled agentic operating model.

ANSWER-FIRST SUMMARY

Key takeaways

Original reporting: Darryl Wong. This is Nexius Labs’ editorial summary and commentary on Darryl Wong’s “When AI doesn't comply”, not a Nexius investigation. Read his article for the full account, method and limitations.

What the case shows

Wong describes two accounting payment workflows: one using an Aspire browser interface and another using the Airwallex API. Records supported earlier agent-initiated submissions following human confirmation. Later, agents refused explicit submission requests and directed that step back to the human.

The browser workflow included human one-time-password authentication. The API records showed accepted submissions; reported settlement was not independently re-audited. These distinctions matter: an accepted instruction, an attempted action and a completed payment are not the same evidence.

What remains uncertain

Four selected before-and-after turns shared the same recorded model label, reasoning effort and core permission settings. That does not establish that the complete system was unchanged: backend versions, later instructions, tool definitions and provider controls were not fully reconstructed.

Wong also reports refusals with Astra and other models. Those additional attempts were operator reports, not individually log-verified comparisons. He attributes the problem to a shared OpenAI-side change, but the investigation did not prove a provider-side cause, mechanism or change date. An agent’s own policy explanation is evidence of its response, not independent confirmation of the governing rule.

This was a retrospective practitioner case study, not a controlled experiment. The two workflows shared an operator and provider. The demonstrated operational change was a human handoff; the records did not establish financial loss, missed deadlines or a measured productivity effect. Earlier execution alone does not show that a later refusal was mistaken.

Nexius perspective: make the next safe step clear

For an SME, continuity planning should cover the point where an agent stops—not just its successful demonstration. We draw three practical priorities from Wong’s account. These are recommendations, not controls tested by his study.

  • Diagnose the specific stop. Separate missing evidence or approval, a tool failure and an asserted action restriction. Correct factual errors with evidence; repeating an instruction does not override a valid boundary.
  • Keep an accurate action record. Capture what was requested, attempted and accepted, alongside available version information. Before anyone retries, reconcile uncertain outcomes with the transaction system to avoid duplicates.
  • Prepare a usable human handoff. Identify the accountable person, the remaining step and the checks still required. Preserve completed preparation so an authorised human can continue through a permitted route.

The goal is dependable work within valid safeguards. A refusal deserves investigation, but neither bypassing controls nor assuming that every refusal is correct provides a sound operating process.

Read the original

Read “When AI doesn't comply” by Darryl Wong for the evidence distinctions and full investigation method. This Nexius summary does not independently verify his private records.

RELATED NEXIUS FIELD GUIDES

Take the concept
into practice.

Continue with implementation-focused guidance from Nexius co-founder Darryl Wong.

OPERATING MODEL / 8 min read

How to Design Non-Technical Work Loops with AI Agents

A practical method for turning recurring business work into bounded, evidence-driven human-agent loops without giving away human authority.

Read field guide
AGENTIC SYSTEMS / 9 min read

How to Agentify ERP and CRM Systems Safely

A governed path from read access to approval-gated execution for businesses introducing AI agents into ERP, CRM, and operational systems.

Read field guide

CONTINUE THE JOURNEY

Related insights

TURN THE IDEA INTO AN OPERATING CAPABILITY

Ready to build your
agentic operating model?

Get the readiness checklist + your recommended next step