Salesforce and Nvidia Ship Koa: an Open-Weight Reasoning Model Built for Enterprise Work, Not Benchmarks
Salesforce and Nvidia have released Koa, an open-weight reasoning model built on Nvidia's Nemotron architecture specifically post-trained for enterprise sales and support workflows.
Announced at the Dreamforce conference, Koa represents Salesforce's first proprietary reasoning model, developed in collaboration with Nvidia. The system leverages Nvidia's open-weight Nemotron as a pre-trained base, addressing a previous gap in available sovereign American models with clear data provenance. Unlike frontier models from proprietary labs that encourage direct data uploads, Koa is designed to operate within Salesforce's Agentforce platform while adhering to strict customer data security requirements. The model has not ingested any actual customer data; instead, the teams crafted synthetic data simulating customer service environments, ranging from irate callers to sales professionals closing deals, to handle specific work tasks rather than abstract benchmarks.
The architectural focus of Koa prioritizes token efficiency and latency over general-purpose capability. According to Nvidia's VP of Generative AI Software for Enterprise, the model utilizes a unique inference architecture optimized for the "trifecta" of sovereign AI, time to first token, and efficient reasoning. This design aims to reduce AI spending by consuming fewer tokens to execute the same multi-step tasks that previously required routing through external frontier models like Claude or ChatGPT via Agentforce's AI gateway. While Koa handles these rote and reasoning-heavy internal tasks, Salesforce maintains its partnership with Anthropic through a new initiative called Claudeforce, which allows enterprises to use Claude as an interface while keeping data secured within Salesforce's infrastructure.
This release signals a divergence between enterprise AI needs and the offerings of frontier labs. Jayesh Govindarajan, EVP of Salesforce AI, noted that prior to Nemotron, no state-of-the-art, sovereign pre-trained model existed that met their criteria for data lineage, contrasting it with models like Qwen where training data origins remain unclear. By providing an open-weight alternative that follows embedded customer security protocols, Salesforce enables automatic routing of requests based on need, allowing agents to resolve long-running tasks locally without exposing sensitive information to external providers. The model serves as a direct alternative to closed systems for customers building agents to schedule appointments or answer service questions, shifting the dependency from external API calls to locally controlled, task-specific reasoning.