The Salesforce Koa reasoning model announced at Dreamforce on Tuesday was post-trained from Nvidia’s Nemotron 3 Super on a synthetic corpus modeled on 27 years of CRM deployments, and the detail that actually matters sits a few paragraphs into Nvidia’s own write-up. Salesforce holds the weights. Salesforce runs the post-training and the inference inside its own trust boundary. And per Rohan Kumar, the company’s president of platform and engineering, not a single byte of customer data was used.
What got clipped and reposted instead was Jensen Huang wandering into the audience at Moscone with a microphone, Marc Benioff trailing him and telling the room they’re all following him. Fine. It’s a keynote, and Huang has earned the victory lap.
I spend a chunk of most weeks inside a CRM for the day job, and that ownership sentence is the whole pitch. The quiet worry about AI bolted onto enterprise SaaS was never that it would hallucinate an opportunity stage. It was whether the pipeline was quietly becoming somebody else’s training corpus. Salesforce answered that with architecture rather than a paragraph in a data processing agreement, and you can audit where weights live in a way you simply cannot audit a contractual promise.
How Koa was actually built
Nemotron 3 Super is an open-weight model, and that licensing fact makes every other claim in this announcement possible. Nvidia spent last summer fronting a 25-signatory letter to Washington arguing for open weights while funding several of the signatories through its $26 billion Nemotron Coalition, and I wrote at the time that the strategy was cleaner than the movement framing suggested: every enterprise that self-hosts is a customer with no pricing leverage over Nvidia, whereas every closed API is one enormous customer with plenty. Koa is that thesis shipping as a product feature at the largest CRM vendor on earth.
The post-training ran as supervised fine-tuning followed by reinforcement learning, using NeMo RL, NeMo Gym and NeMo AutoModel. Gym is the piece worth understanding, because it’s where the difference between a chat model and an agent model gets made. You’re not rewarding the model for producing pleasant prose about a sales opportunity. You’re dropping it into a simulated environment with tools and a schema, letting it attempt a multistep workflow, and scoring whether the workflow completed correctly. AutoModel handles the fine-tuning plumbing across model formats, RL runs the loop at scale, and the output is a model whose competence is measured in completed actions rather than token quality.
The corpus is where Salesforce got clever. Rather than train on real deployments, which would blow up the entire privacy claim, they generated synthetic scenarios pairing a persona with a task, spanning more than 14 industries including financial services, manufacturing, healthcare and travel. The tasks are the boring ones: generating a lead, qualifying an opportunity, resolving a service case. Privacy is the stated reason for going synthetic, and it’s a real one, but there’s a second benefit nobody puts in the press release. Real CRM data is overwhelmingly the happy path. Synthetic generation lets you manufacture the ugly cases on demand: the duplicate contact attached to a half-filled account record with the wrong currency on the opportunity, and label them correctly, which is precisely the data you cannot get enough of from production logs.
This is why a mid-size specialist makes sense here, and it took me a while to stop thinking of it as a downgrade from frontier models. The failure mode in a CRM agent is almost never insufficient intelligence. It’s picking the wrong tool, writing to the wrong field, or silently dropping a required parameter. Those failures compound viciously across multistep work: a model that’s 95% reliable per action finishes an eight-step workflow correctly about two-thirds of the time, and getting per-action reliability to 98% takes you to roughly 85% end to end. That’s the arithmetic that makes Salesforce’s headline claim, matching or beating leading models on CRM actions with three times fewer errors on its CRM Bench suite, a more useful number than any general reasoning score you’ll see quoted this quarter.
Assuming you trust it, which I don’t yet. It’s Salesforce’s benchmark, Salesforce’s task selection, unnamed competitors, and no published definition of what counts as an error. Three times fewer than what, exactly, and measured how? I want somebody outside the building to run a frontier model against the same suite before anyone treats that multiplier as load-bearing. In fairness, a number with holes in it still beats what ServiceNow did when it announced Apriel 2.0 with no benchmarks at all.
One architectural consequence deserves more attention than it got. If no customer data touches training or inference, the model never learns your organization. All the specificity arrives at runtime, through the context and permissions layer, which means the weights stay generic and your data stays yours. Good for compliance, and quietly strategic: a generic model is genuinely swappable, while the governed context layer underneath it is not. Salesforce built a thing it can replace and kept the thing it cannot. Read the rest of the Dreamforce announcements with that in mind, and the priorities get obvious fast.
ServiceNow ran this play first
Here’s the part that undercuts the “first CRM reasoning model” framing. ServiceNow and Nvidia shipped Apriel Nemotron 15B at Knowledge 2025, built with NeMo, Nvidia’s Llama Nemotron post-training dataset and ServiceNow’s own workflow data, trained on DGX Cloud. Same structure, same toolchain, same argument about lower latency and lower inference cost from a small domain model. They followed it with Apriel 2.0, announced at GTC for Q1 2026 availability, explicitly aimed at autonomous agents in regulated industries and wired into Nvidia’s AI Factory reference designs.
So Salesforce’s debut is roughly ServiceNow’s third iteration, and what I find more interesting than either company is the template. Nvidia has a repeatable kit now: take the platform vendor’s proprietary workflow knowledge, post-train Nemotron on it, let the vendor keep the weights, and collect on the compute underneath. Every enterprise software company with a decade of process data sitting in a warehouse is a candidate. SAP, Workday, HubSpot, Zendesk, pick your vertical. I’d be surprised if we don’t see at least two more of these before the middle of next year, and they’ll all sound like this one.
Microsoft attacks from the opposite direction, with integration depth on top of frontier closed models it licenses rather than owns. That depth is a genuine moat and also the exact thing regulators keep catching. When I went through AXA’s Copilot rollout and the sovereignty gateway it walked around, the recurring problem wasn’t model quality; it was the ambient semantic index and the fact that CLOUD Act exposure attaches at the corporate level regardless of which data center the bits sit in. A model whose weights live inside the vendor’s trust boundary doesn’t fix that for a European buyer, but it’s a materially different conversation from “our index calls out to an American frontier lab.”
And the frontier labs themselves were in the room, which is the tell. Salesforce announced Claudeforce at the same event. TechCrunch called Koa everything the AI labs should fear, which is a touch dramatic, though the underlying point holds: Salesforce isn’t trying to replace Anthropic or OpenAI; it’s demoting them into one entry in a picker. Koa ships as a customer-selectable model inside Agentforce. That phrase is doing an enormous amount of work.
What it costs to own a model you didn’t train
Strip away the announcements, and there are only a few ways a company can buy enterprise AI right now. You can rent tokens from a frontier lab and accept the dependency. You can self-host open weights and own the whole stack, staffing included. Or your platform vendor can do the second one on your behalf and bill it as a feature, which is what Koa is, and which is the option that actually fits the enormous middle of the market.
The numbers behind that middle are unforgiving, and I’ve run them before. Self-hosting starts breaking even against frontier API pricing somewhere around 100 million tokens a month and only wins decisively past half a billion. Below roughly 20 million, it’s nowhere close. The dominant line item is never the GPUs; it’s the $700,000 to $1.4 million a year in engineering, monitoring, and incident response that a serious self-hosted deployment quietly requires. Almost no mid-market company clears that bar, which is why the sovereignty conversation in most boardrooms dies somewhere between the ambition and the headcount request.
Koa sells you the outcome of that build without the build. You get residency, a controlled trust boundary, and a model tuned to your actual workflows, at the cost of the ownership being Salesforce’s rather than yours. That’s a reasonable trade for a lot of companies and a bad one for a few. If your differentiation lives in how you run your processes, renting a model trained on the aggregate of everyone else’s processes is a strange place to plant a flag.
For anyone doing procurement on this, the questions that matter aren’t about benchmarks. Find out how the model is priced against your consumption terms, because a vendor-owned small model has radically better unit economics than a frontier API call, and you should be capturing some of that margin rather than all of it flowing upstream. Find out what “customer-selectable” means when you want to select something else eighteen months in, and whether the agents you build are portable across that choice. Ask what happens to the context layer if you leave, since that’s the expensive part. And if you’re in the EU, note that general availability is winter 2026 in US regions, so the question is academic for a while yet.
The most revealing deployment target isn’t commercial at all. Salesforce is pushing Nemotron-based models into Missionforce, its defense work, aimed at private clouds and air-gapped networks, with post-trained Nvidia models reaching selected customers in October. That’s the logical endpoint of this architecture. Once a competent model can run entirely inside a trust boundary you control, it can run inside a classified one, and the addressable market for enterprise AI stops ending at the edge of the public cloud.
Where I think this lands
Pilots start in October with Formula 1, UChicago Medicine, Baxter Credit Union, 1-800Accountant, Engine and Xero. General availability is penciled in for winter, US first. The number I care about won’t exist until well into next year, when somebody in one of those pilots publishes an action-level error rate from production rather than from a benchmark, and my guess is it lands meaningfully worse than the lab figure and still good enough to justify the switch. That’s usually how this goes.
Between now and then I’d expect the frontier labs to respond by climbing into the workflow layer themselves rather than defending the model layer, because defending the model layer against free open weights post-trained by your own customers is not a fight with a good ending. Anthropic and OpenAI both have obvious paths into enterprise workflow, and both have obvious reasons to resent being a dropdown.
None of which threatens Nvidia, and this is the bit I keep landing on no matter which direction the app layer moves. The CUDA position compounds rather than erodes precisely because open weights multiply the number of places CUDA runs. Nvidia gives away the base model, supplies the post-training stack, sells the silicon, helps arrange half a trillion dollars of financing so its customers can buy the silicon, and reportedly bought the repository where the weights are distributed. Openness here is a distribution strategy with a hardware invoice attached, and it costs Nvidia nothing to let Salesforce own weights that only ever execute on Nvidia’s floor.
What stays with me is that dropdown. If the reasoning model becomes a component your CRM vendor swaps at will, then a hundred billion dollars of frontier training spend bought a part, not a platform. I don’t believe that’s true for genuinely hard reasoning work, and I’d bet against anyone who says otherwise. For updating an opportunity and routing a case, though, it already looks true.
Sources
- NVIDIA Blog, ‘Now We Can Know Everything and Do Anything,’ Jensen Huang Says at Dreamforce, September 15, 2026
- TechCrunch, Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear, September 15, 2026
- Salesforce Ben, Salesforce’s First CRM Reasoning Model ‘Koa’ Is Revealed at Dreamforce ’26
- ServiceNow Newsroom, ServiceNow Unites Intelligent Workflows and Open Models with NVIDIA Technologies, on Apriel 2.0
- CIO, ServiceNow’s Apriel 2.0 promises smarter AI with less hardware, but offers no benchmarks