When AI Agents Help and When They Hurt
The line between agentic value and agentic failure is real, invisible, and unmarked || Edition 33
This post is part 15 of the Agentic AI Series — a multi-part exploration of how autonomous systems are reshaping enterprise architecture, governance, and security.
Klarna, Feb. 2024. The Swedish fintech deployed an OpenAI-powered customer service assistant. The company said it handled 2.3 million conversations in its first month, the equivalent of 700 full-time agents, and was projected to add $40 million in profit.
Klarna, May 2025. CEO Sebastian Siemiatkowski told Bloomberg that cost had dominated the evaluation and quality had collapsed. Customers described the responses as generic and repetitive. It recited documentation and transferred users to human support, acting more as a filter than a resolution engine.
Klarna, mid-2025. The company launched a hiring push for human agents. The CEO revised his position: human connection would define the VIP experience, while AI customer service would remain the cheap option.
One of the most visible AI workforce replacement stories in enterprise reversed in under eighteen months.
The Short Version
AI value is uneven. The Harvard and BCG study found that the same model that improved speed, volume, and quality on some consulting tasks reduced accuracy on others.
The jagged technological frontier is the boundary between where AI can be trusted and where it must be checked. That boundary does not move smoothly across tasks, and teams get no visible signal for where reliable performance ends and degradation begins.
Frontier-mapping is an economic discipline. Klarna, McDonald’s, and Air Canada each deployed AI into workflows that looked routine, then absorbed the cost of discovering in production that the systems could not handle the real complexity of the job.
Human involvement remains a structural requirement in the strongest operating models. Centaurs divide labor deliberately. Cyborgs stay engaged throughout the task. Both preserve judgment inside the workflow and outperform full delegation.
Agentic systems amplify the cost of being wrong. In a human-in-the-loop workflow, a wrong answer can be caught or revised. In an agentic workflow without explicit human intervention, that same error can trigger an action before anyone intervenes.
The Jagged Frontier
AI improves some tasks and degrades others, a pattern demonstrated in a Harvard and BCG study published in Organization Science.
The researchers call this uneven boundary the jagged technological frontier. The frontier is the boundary between where AI can be trusted and where it must be checked.

Inside the frontier, AI improved speed, volume, and quality. Outside it, performance degraded on tasks that required judgment the model did not have. GPT-4 produced persuasive but wrong answers, and people followed them. The leader author described this as falling asleep at the wheel.
Two additional findings matter for leaders.
AI compressed performance differences by helping lower performers more than higher performers.
It made outputs more similar, which raises the risk of strategic sameness when many teams rely on the same models for the same work.
The lesson: Value from AI is task-specific. The boundary is invisible, and it has to be mapped before deployment.
Agentic Systems Raise the Cost of the Jagged Frontier
Autonomy turns frontier errors into operational consequences. The BCG and Harvard study measured human-AI collaboration, where people could accept, reject, or revise outputs before they mattered. Agentic systems compress or remove that checkpoint.
Klarna’s bot delivered the answer directly to the customer. The system moved beyond the point where its output could be trusted and acted before a human could intervene.
Key Takeaway: Agentic systems raise the cost of crossing that boundary because unreliable outputs can become executed consequences.
How Strong Teams Actually Work With AI
The strongest deployments keep human judgment inside the workflow by design. The Harvard and BCG study surfaced two collaboration patterns that consistently created value.
Centaurs place judgment at the handoff and delegate only what fits the frontier. Cyborgs keep judgment active throughout the task as human and model work in continuous interaction. Full delegation removes the checkpoint while the frontier still exists. That is where failures compound.
The lesson: Strong systems decide where judgment must remain embedded in the workflow and design for it explicitly.
From Field Lessons to Operating Decisions
Lessons from the Field
Real deployments have already shown where full delegation breaks down.
Tacit knowledge matters more than it looks.
McDonald’s deployed IBM’s voice AI across more than 100 U.S. drive-thrus and pulled it back after repeated failures. The workflow depended on accent variation, background noise, order modifications, and implicit customer context that experienced workers handle in real time. The system could only act on what had been made explicit. Many frontier failures begin where tacit knowledge exceeds codified knowledge.
Organizations remain accountable for AI-delivered information.
Air Canada’s customer chatbot fabricated a bereavement discount policy and told a grieving customer he could retroactively apply for it. The dispute escalated into a legal claim after Air Canada refused to honor the information its chatbot had provided. The tribunal rejected Air Canada’s argument that the chatbot was a separate legal entity responsible for its own actions. Organizations are accountable for what their AI tools say and do.
Guardrails require active oversight in production.
DPD’s chatbot went rogue after a routine system update disrupted its controls. It swore at customers and called DPD the worst delivery firm in the world. The incident went viral. Guardrails need continuous monitoring, and systems need human escalation paths when outputs drift beyond policy or tone.
Each of these failures came from deploying before the workflow had been proven. The organizations succeeding with AI validate value before they scale it.
These failures point to a practical question. How should organizations choose use cases, design workflows, and govern systems before those same patterns show up in production?
The Decisions That Determine Outcomes
Selecting use cases
Choose workflows that fit the frontier. Do the inputs, boundaries, and outcomes make the work observable and testable?
Separate assistance from autonomy. Where does AI help, and where must judgment remain inside the workflow?
Building systems
Test the entire workflow. Does performance improve end to end, or only at a single step?
Decide where judgment lives. Where should a human evaluate, redirect, or override before the system moves forward?
Measure system shift. Does the workflow get better, or does one improvement create downstream fragility?
Managing projects
Measure impact at the process level. Did the system become more accurate, resilient, and effective?
Make frontier-mapping explicit. Where does AI help, where does it hurt, and where does judgment change the outcome?
Watch for convergence. Does improved consistency come at the cost of differentiation?
Governing deployment
Require workflow evidence. Where has this system been proven to operate reliably within the frontier?
Assign accountability clearly. If the system speaks or acts, who owns the consequence?
Sequence workforce changes with learning. What do you still not know about where value is created or lost?
In Closing
AI value is a property of the workflow, the context, and the placement of judgment inside the system.
The real question is whether the work has been chosen well, whether the frontier has been tested, and whether the workflow can hold when the system reaches its limits.
The teams that get this right will look disciplined before they look ambitious. They will choose narrower use cases, keep humans where judgment matters, and treat value as something to validate before it is scaled.
Diagnostic question for your next meeting: What empirical evidence exists before this system is allowed to scale?
The next post shifts from where agentic systems create value to what it costs to keep them governable, observable, and resilient in production.



