← live missions
analysis 2026-09-19

Does agentforce work? The agentic customer-service record before you buy, 2026

Salesforce sells Agentforce as agents that resolve cases and hand off to humans only when needed. The last big autonomous customer-service deployments say the dial you set matters more than the demo.

You are evaluating Agentforce because Salesforce is telling your board that AI agents can run customer service on their own. The pitch is specific: agents that “resolve cases,” operate “24/7,” and “hand off to humans only when needed” [source]. The question your team is actually asking is narrower and harder. Does it work.

Nobody outside Salesforce can answer that with a measured outcome yet, and TIN will not pretend otherwise. There is no independently verified Agentforce customer-service result in our corpus. What we do have is the verified record of the last big autonomous customer-service deployments, and that record answers a better question than “does the product work.” It tells you which setting on the autonomy dial survives contact with real customers.

What Agentforce is selling

Agentforce is Salesforce’s platform for autonomous customer-service agents that use large language models to “reason through decisions autonomously” and take action, escalating to a person only when a request exceeds their scope [source]. The load-bearing word is “autonomously,” and the load-bearing qualifier is “only when needed.”

That qualifier is a dial, not a fixed position. You choose how much the agent closes on its own and how much a human sees before it reaches the customer. The record is about where that dial was set, not about whose logo was on the software.

Does agentforce work: what the closest cases show

There is no verified Agentforce figure to quote, so the honest answer is to show the two deployments that sit on either end of the same dial Salesforce is asking you to set.

Klarna set it toward full autonomy and toward cost. Its OpenAI-powered assistant handled two-thirds of chats in month one, February 2024, with a modeled equivalence to 700 full-time agents and a 25% drop in repeat inquiries [source]. The figures were real and first-party. Then in 2025 CEO Sebastian Siemiatkowski said the cost-first push had produced “lower quality,” and Klarna began re-recruiting human agents ([source]).

Octopus Energy set the same dial the other way. Its Kraken “Magic Ink” tool drafted email replies, with a human agent reviewing and sending each one. By end-April 2023 CEO Greg Jackson said it handled 34% of customer queries, the work of about 250 people in the UK, and the company stated there would be no job cuts [source]. Kraken’s own case study confirms a person reviews and sends every reply ([source]). Octopus later announced 4,000 new roles.

Same year, same department, same “work of X people” rhetoric. One deployment reversed and one held. The difference was not the model. It was the dial.

What “hand off to humans only when needed” costs when it fails

Salesforce’s escalation qualifier assumes the agent knows when it is wrong. Sometimes it does not, and the record has a price for that.

In Moffatt v. Air Canada a BC tribunal found the airline had negligently misrepresented its bereavement-fare policy through the chatbot on its own website, and ordered C$812.02 [source]. The airline argued the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal rejected that outright, holding the company responsible for all information on its site whether it comes “from a static page or a chatbot” ([source]).

The dollar figure is small. The principle is not. An autonomous agent that acts “on your behalf” acts as you, in law, and buying it from a vendor does not move that liability to the vendor. The more the agent closes without a human, the more of that exposure you are carrying.

The proof

TIN verified each of the three case files behind this post against the public record before it was written. None of them is an Agentforce deployment, and this post claims no Agentforce result. They are the closest verified analogues to the setting Salesforce is asking you to choose.

How the two dial settings compare

SettingDeploymentHuman in loopHeadline figureOutcome to date
Autonomy-first, cost-firstKlarnaMinimalTwo-thirds of chats, month oneWalked back, rehiring humans
Draft, human sendsOctopus EnergyHuman sends every reply34% of queries, work of ~250 peopleHeld, 4,000 new roles
Autonomous, on its own websiteAir CanadaNone on the answerC$812.02 orderedCompany held liable

The bottom line

Do not ask whether Agentforce works. Ask where you are going to set the dial it ships with, because that is the choice the record actually judges. The verified evidence points one way: the deployment that kept a person on the send button held and grew, and the two that let the agent close on its own either reversed on quality or paid for a wrong answer in a tribunal. The product is unproven in public; the setting is not.

Buy the deployment model, not the demo, and price the liability you keep either way.

Sources

  1. 01
    Salesforce, “Agentforce: Create powerful AI agents,” retrieved 2026-09-19. https://www.salesforce.com/agentforce/
  2. 02
    Forbes / Jack Kelly, “Klarna’s AI Assistant Is Doing The Job Of 700 Workers, Company Says,” 2024-03-04. https://www.forbes.com/sites/jackkelly/2024/03/04/klarnas-ai-assistant-is-doing-the-job-of-700-workers-company-says/
  3. 03
    Customer Experience Dive, “Klarna changes its AI tune and again recruits humans for customer service,” 2025-05. https://www.customerexperiencedive.com/news/klarna-reinvests-human-talent-ai-customer-service-buy-now-pay-later/747586/
  4. 04
    City AM, “AI doing the work of over 200 people at Octopus, chief executive says,” 2023-05-08. https://www.cityam.com/ai-doing-the-work-of-over-200-people-at-octopus-chief-executive-says/
  5. 05
    techUK, “Case study: Kraken Tech’s generative AI tool for customer service,” 2023. https://www.techuk.org/resource/case-study-kraken-tech-s-generative-ai-tool-for-customer-service.html
  6. 06
    BC Civil Resolution Tribunal (CanLII), “Moffatt v. Air Canada, 2024 BCCRT 149,” 2024-02-14. https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
  7. 07
    Ars Technica / Ashley Belanger, “Air Canada must honor refund policy invented by airline’s chatbot,” 2024-02-16. https://arstechnica.com/tech-policy/2024/02/air-canada-must-honor-refund-policy-invented-by-airlines-chatbot/

Questions

Does agentforce work?

There is no independently verified Agentforce customer-service outcome in TIN's corpus yet, so nobody outside Salesforce can answer that with a measured result. What the record does show is that the autonomy setting matters more than the product: the last big cost-first autonomous deployment, Klarna, walked back to human agents, while the human-in-the-loop deployment, Octopus Energy, held and kept hiring.

Is Agentforce fully autonomous or does it keep a human in the loop?

Salesforce markets Agentforce as autonomous support that will 'hand off to humans only when needed.' That is a dial, not a fixed setting, and the record says where you set it is the whole decision. The deployment that put a human on every send held up; the one that pushed cost-first autonomy reversed.

Who is liable if an autonomous agent gives a customer wrong information?

The company is. In Moffatt v. Air Canada a tribunal held the airline responsible for its website chatbot's bad advice and rejected the argument that the bot was a separate legal entity, ordering C$812.02. Buying an autonomous agent does not move that liability to the vendor.

This is analysis, not a verified outcome. It carries no verification badge and never will. The proof lives in the case files, where every figure is checked against the public record and the method is printed on the page.