Skip to news

Cisco Moves Enterprise AI Back On Premises With Splunk

Cisco’s new AI POD for Splunk turns data sovereignty, cost control, and agent oversight into a single enterprise infrastructure bet.

By THE COLDAI TIMES deskPublished 5 min read1,030 words

Cisco is betting that the next phase of enterprise AI will not be defined by sending more corporate data to hyperscale clouds. Instead, it is bringing AI workloads back into the data centers, private clouds, and air-gapped environments where many companies still keep their most sensitive operational information.

At Splunk’s .conf26 conference in Denver on September 15, Cisco introduced Cisco AI POD for Splunk, a prevalidated infrastructure package that combines Cisco systems, NVIDIA accelerated computing, Kubernetes-based software, and Splunk’s enterprise AI capabilities. The product is designed to let customers run AI Assistant, model-powered investigations, and future agent-building tools inside their own environments rather than moving security and machine data to an external service. (newsroom.cisco.com)

The announcement is more consequential than a conventional product launch because it addresses three constraints that have begun to collide in enterprise AI: data sovereignty, operational risk, and cost. Companies may want autonomous systems to investigate incidents or operate business workflows, but many cannot place logs, credentials, network telemetry, industrial data, or regulated records into a third-party cloud. Cisco’s answer is to package the compute, runtime, and observability layers together so organizations can deploy AI without redesigning their security architecture from scratch.

What changed

The new AI POD is available for Splunk Enterprise customers running workloads in on-premises, private-cloud, and air-gapped settings. Cisco says the system supports Splunk AI Assistant and a selection of hosted models, including Google’s Gemma 4 and OpenAI’s GPT-OSS 20B, with NVIDIA Nemotron models expected in the coming months. That model choice matters: customers are not being asked to accept one universal model or one external provider. They can select models according to the workload, governance requirements, and performance profile. (newsroom.cisco.com)

Cisco also introduced Tokenomics within Splunk Agent Observability. The tool is intended to show how much AI agents and coding assistants are being used, what they cost, and whether that spending produces measurable business value. Separate observability updates are designed to connect agent behavior with application, network, and infrastructure performance.

That combination reflects a shift in enterprise buying criteria. The question is no longer simply whether a model can summarize an alert or generate code. IT leaders increasingly need to know which agent acted, what data it accessed, how many model calls it made, what those calls cost, and whether its activity caused an operational problem. Network World described the release as a package for monitoring agent behavior, controlling token costs, and correlating AI activity across applications and networks. (networkworld.com)

Cisco is also formalizing a multiyear agreement with AWS to develop security products aimed at AI-driven attacks. The arrangement shows that the company is not treating on-premises AI as a rejection of cloud infrastructure. Rather, Cisco is positioning itself as a layer between multiple deployment models: private infrastructure for sensitive workloads, public cloud where elasticity is useful, and security telemetry that can span both.

Why it matters

The strategic importance lies in infrastructure placement. For the past several years, the default enterprise AI pattern has been to centralize data in cloud platforms and call remote models through application programming interfaces. That model remains attractive for experimentation, but it is often difficult to reconcile with national data rules, sector-specific regulation, internal security policies, or the economics of moving large volumes of machine data.

Cisco’s launch suggests that the market is entering a hybrid phase in which inference location becomes a board-level decision. A bank may want an agent to investigate transactions without exporting raw customer data. A hospital may need clinical or operational systems to remain isolated. A defense contractor may require an air-gapped deployment. A manufacturer may prefer to keep factory telemetry near the equipment because latency and intellectual-property concerns make cloud transfer undesirable.

Running locally does not automatically make those systems safe. It can reduce exposure to external data transfers, but it also shifts responsibility to the customer. Organizations must maintain hardware, patch model-serving infrastructure, secure model files, control agent permissions, and monitor internal misuse. An on-premises agent with broad access to security tools can still create damage faster than a human analyst, particularly if the system is connected to ticketing, identity, endpoint, or network controls.

That is why the observability layer may be as important as the AI hardware. The recent acceleration of agentic systems has exposed a basic governance problem: conventional application monitoring was built to track deterministic services, not systems that can choose tools, repeat actions, delegate tasks, and consume variable amounts of compute. Token and behavior monitoring are early attempts to make those systems legible to finance and security teams.

The deeper bet is that enterprise AI will become a control-plane market. Cisco is not competing only on model quality. It is trying to own the operational environment around models: the servers, networking, security data, observability, and policy controls required to make agents acceptable in production. NVIDIA supplies much of the accelerated compute; Splunk supplies the data and security context; Cisco supplies the integration and enterprise channel.

What remains uncertain

The announcement does not establish how widely customers will adopt the platform, what a typical deployment will cost, or how much performance customers will sacrifice compared with hyperscale inference. Cisco says the AI POD is generally available, but availability is not the same as demonstrated production scale. The practical economics will depend on model size, utilization, GPU capacity, support contracts, and whether customers can keep the systems busy enough to justify the capital expense.

There are also unresolved questions about openness. Cisco lists several supported models, but enterprises will want to know how quickly new models can be added, whether customers can bring their own models, and how deployment policies differ between proprietary and open-weight systems. They will also need evidence that the monitoring tools can detect subtle agent failures rather than merely report usage and latency.

The launch therefore marks a direction, not a settled market outcome. Cisco is responding to a clear enterprise hesitation: AI cannot scale if the data required to operate it cannot legally, economically, or safely leave the building. But putting AI back on premises will not eliminate risk. It will relocate the risk into infrastructure that customers must now be capable of operating, auditing, and defending themselves.

Related stories