Ten AI agents can still give you one opinion.
Anthropic's August 2026 research on emerging multi-agent systems gives operators a useful warning. When agents share the same model, context, scaffolding, and incentives, they often make the same choices. In one experiment, 18 of 30 agents created the same branch name. In another, more than half independently chose one of two similar project types.
That is not diversity. It is one blind spot multiplied across more workers.
For SMEs moving from chat to execution, this changes how an AI workforce should be designed. More agents can increase throughput, but they do not automatically increase judgment. A reliable digital coworker team needs independent evidence lanes, a deliberate dissent role, and one accountable human who integrates the result.
The Agent Count Trap
Many teams treat multi-agent design as a headcount question: if one agent helps, five agents must be better. They assign labels such as researcher, analyst, critic, and manager, then give every role the same brief and the same source pack.
The system looks organised. Underneath, the agents may still share the same assumptions, miss the same exception, and reinforce the same confident answer.
Anthropic describes current agents as relatively low variance. When model, context, and scaffolding are similar, agents can take surprisingly similar actions even when many alternatives exist. What would be an isolated mistake in a single-agent workflow can become a systemic mistake when every agent repeats it at machine speed.
For an SME, the failure will not always look dramatic. It can appear as five sales agents qualifying the same poor-fit leads, three finance agents missing the same absent record, or a procurement team summarising different vendor proposals through the same biased criteria.
The New Agent Boss Job: Manufacture Useful Disagreement
The Agent Boss is not the person who watches the most agent activity. It is the domain owner who designs how an AI team reaches a decision.
That means separating three jobs:
- Build the strongest supported recommendation.
- Search for evidence that could overturn it.
- Integrate the disagreement into an accountable decision.
This is designed dissent. It does not mean asking agents to argue for entertainment. It means making one lane responsible for finding the missing fact, policy conflict, edge case, or alternative explanation that the main workflow could ignore.
A confident consensus is not evidence of correctness when every participant was built from the same starting point.
A Three-Lane Pattern for SME Agent Teams
1. Builder lane
The builder prepares the recommendation from approved operational data. It should return the proposed action, source evidence, assumptions, missing information, and confidence limits.
Example: a receivables agent prioritises overdue accounts and drafts the next action using invoice age, dispute status, customer history, and credit policy.
2. Dissenter lane
The dissenter does not rewrite the same recommendation in different words. It receives different evidence and a different acceptance test. Its job is to find what would make the recommendation unsafe, incomplete, or wrong.
For receivables, that might include unresolved service complaints, credit notes, disputed deliveries, strategic-account status, or a policy exception that is absent from the main ledger view.
3. Integrator lane
The integrator compares both evidence packets. It preserves material disagreement, identifies missing data, and routes the decision according to written thresholds.
The integrator may prepare a concise decision brief. It should not erase a minority view behind one synthetic confidence score. A named finance or operations owner remains accountable for consequential action.
Different Roles Need Different Evidence
Role labels are cosmetic unless they change the work.
If the builder and dissenter use the same prompt, same model, same documents, and same success metric, the SME has created duplicate labour. To create useful independence, vary at least some of the following:
- Source access: operational records for the builder; policies, complaints, failed cases, and exception history for the dissenter.
- Question: “What should we do?” for the builder; “What evidence would reverse this?” for the dissenter.
- Model or method: where practical, use different models, retrieval paths, or deterministic checks for critical facts.
- Success metric: quality of the recommendation for one lane; quality of discovered exceptions for another.
- Permissions: evidence agents may read; execution permissions remain narrow and separately controlled.
The aim is not artificial disagreement. The aim is to reduce correlated failure.
Classify the Work Before Adding Agents
Anthropic's experiments also show that multi-agent systems perform differently depending on the work. Independent vulnerability-search approaches and a coordinated swarm found largely complementary results, with only 12 vulnerabilities in common. But when agents depended heavily on one another during game development, coordination became harder and the end products remained poor despite role prompts and CEO hierarchies.
SMEs should classify each workflow before choosing an agent pattern:
- Independent work: parallel research, evidence collection, document review, scenario generation, or separate quality checks. Multiple agents can be useful.
- Dependent work: tasks where one output becomes another agent's input and errors compound across handoffs. Use fewer agents, tighter interfaces, and stronger integration tests.
- Integrator-owned work: decisions involving policy, money, customers, people, or business trade-offs. Name the human owner before increasing automation.
A swarm is not a universal upgrade. Use agent count only where the work can be separated cleanly or where independent evidence has real value.
Measure the Team, Not Just Each Agent
Individual agents can pass their tests while the overall workflow still fails. Team-level telemetry should show whether the system produces independent evidence, preserves important disagreement, and avoids machine-speed bureaucracy.
Track:
- evidence overlap: how much of the source set is genuinely independent;
- disagreement rate: whether supported alternative conclusions ever appear;
- reversal quality: whether new evidence changes the recommendation appropriately;
- integration loss: whether caveats disappear from the final brief;
- human correction rate: what the owner changes and why;
- queue health: duplicate requests, abandoned handoffs, review age, and accepted outcomes.
One Anthropic queue experiment is a sharp reminder. Agents generated 2.4 million job requests for only 117 accepted jobs. Every worker could look active while the shared system became worse.
Measure completed business outcomes and review quality, not messages, attempts, or agent activity.
A Practical Pilot for SMEs
Choose one decision-preparation workflow: supplier shortlisting, overdue-account prioritisation, sales qualification, policy-exception review, hiring screening, or customer-escalation triage.
- Run the current single-agent or human process on 20 to 50 historical cases.
- Add a dissenter lane with different evidence and the required question: “What fact would overturn this recommendation?”
- Require both lanes to cite source records and state missing information.
- Have the named domain owner integrate the cases without hiding disagreement.
- Compare exception discovery, human correction, cycle time, review burden, and decision quality.
- Keep execution in shadow mode until the team can fail safely and its decisions can be reconstructed.
This moves an SME from awareness to adoption without pretending that a larger agent team is automatically a mature one.
Orchestrate, Don't Duplicate
The value of an AI workforce is not the number of agents on an architecture diagram.
The value is whether the team can find different evidence, challenge a weak assumption, hand work over cleanly, and help the accountable human make a better decision.
Do not copy one digital coworker ten times and call it judgment. Design the evidence lanes. Assign the dissenter. Name the integrator. Then measure whether the business decision improves.
Sources
- Anthropic: Patterns and problems in emerging multiagent systems
- Hacker News discussion of the Anthropic research
- Anthropic Engineering: How we built our multi-agent research system
Need help designing an SME agent workflow that produces independent evidence instead of duplicated output? Nexius Labs helps teams map, build, test, and operate digital coworker systems with clear ownership, telemetry, human review, and auditability.
Start Here

