At a Glance
- Salesforce launched Agentforce 360 on September 11 with seven named agents. Six are GA. The seventh, Hunter, is the one worth watching.
- Hunter runs on a new long-horizon runtime: memory, durable execution, and dynamic steering that let it pursue a single goal across days or weeks, not a chat session.
- If it works, entire process categories become agent-native. If it doesn’t, three weeks is a long time to be silently going wrong.
Six agents shipped. One quietly changes the game.
When we wrote about the Salesforce-Anthropic partnership last week, the open question was what it would actually ship. On September 11, we got the answer: seven named agents in Agentforce 360.
Salesforce introduced Casey, Paige, Carter, Marshall, Piper, and Fin – service, IT/HR, commerce, supply chain, inbound sales, and customer experience (Fin came in with the Intercom acquisition on September 10). Six are GA. All six are polished, fast, single-session responders. Useful, but not new in kind.
The Seventh – Hunter is different. It’s the first agent Salesforce has shipped on a long-horizon runtime, designed to preserve state, hold context, and pursue a goal across days or weeks, using memory, durable execution, and dynamic steering that lets humans nudge it mid-flight without stopping the whole thing. It’s still in pilot, with GA targeted for November.
Why long-horizon is the unlock
Most enterprise processes worth automating aren’t sessions. A renewal cycle. A hiring loop. An outbound sales campaign. An incident investigation. They run for weeks, involve many touchpoints, and demand continuity of context that a single-session agent cannot hold. This is where Generative AI services start moving beyond single-session assistance and into longer-running business processes.
Everything we’ve deployed so far has been Lego blocks, capable agents you snap together with a human orchestrator sitting on top. A long-horizon runtime is the first serious attempt to remove the human as the connective tissue. One agent, one outcome, weeks of runtime, with escalations to humans on exception rather than by default.
Why it’s hard
Getting an agent to behave sensibly across a fifteen-minute chat is hard. Getting it to behave sensibly across three weeks is a different problem:
- State without drift. The agent has to remember what it committed to, and not silently rewrite that memory when new context arrives.
- Judgment about time. Knowing when to act, when to wait, when the situation has changed enough to escalate.
- Failure modes at both ends. Going quiet (stalling) is as damaging as going rogue. Both are hard to detect until the outcome is already off.
- Durable execution. Weeks-long processes need checkpointing, replay, and resumability that most current agent stacks don’t have.
- New QA. You can’t regression-test a three-week agent by replaying a transcript. Evaluation has to move from session-level to outcome-level.
- Cost shape. Agents that run continuously don’t cost like agents that answer questions. Your FinOps model needs a rewrite.
None of this is a reason not to pursue it. But it does mean Hunter’s November GA is a start-of-the-conversation date, not a deploy-to-production date.
Atgeir’s take:-
If you have an outbound motion, or a renewal cycle, or any process that’s naturally weeks long, then get on the Hunter pilot list. This is the direction of travel, and the teams that learn to operate long-horizon agents first will pull ahead sharply.
But plan the human-in-the-loop before you plan the agent. A long-horizon agent needs a review cadence, escalation paths, and outcome metrics that look nothing like what your CRM team owns today. That makes operating strategy as important as agent selection, and increasingly relevant to ai strategy consulting.