Network management · 4 MIN READ

The Next Frontier for AI Fabrics: Scaling Across Networks

AI infrastructure can no longer be planned solely as one vertically integrated cluster. As physical space and available power become limiting factors, scale-across fabrics extend AI compute over long-distance networks—changing architecture, capacity planning, and operational visibility across sites.

The Next Frontier for AI Fabrics: Scaling Across Networks

From Centralized Clusters to Scale-Across Fabrics

AI training workloads and frontier models continue to grow toward trillions of parameters and millions of accelerators. At that scale, adding more equipment to a single location runs into a pragmatic constraint: data center space and power are finite.

The scale-across model addresses this limitation by distributing compute density across geographies. It extends the concepts of local scale-up and scale-out clusters over long-distance network connections, shifting AI infrastructure from a centralized vertical stack toward a horizontal, multi-site fabric.

Scale-across is not simply a larger local cluster. It makes the network between locations part of the AI system’s architecture.

How Network Architecture Changes

In a centralized design, accelerator connectivity, compute resources, and much of the supporting infrastructure can reside within one facility. A distributed fabric introduces dependencies that cross physical locations.

Network teams therefore need to consider:

  • Inter-site connectivity as a core dependency: Long-distance paths are no longer peripheral links when they connect parts of the same AI environment.

  • Geographic placement: Compute must be mapped to locations with suitable space, power, and network capacity.

  • Cross-site failure domains: A site, path, or configuration change can affect resources beyond one local cluster.

  • Configuration consistency: Distributed network devices must operate as one coordinated infrastructure even though they are managed across different facilities.

  • End-to-end topology: Operators need to understand how local scale-up and scale-out domains connect through the wider scale-across fabric.

This changes the architectural question from “How large can this cluster become?” to “How can multiple locations operate as one coherent compute fabric?”

Capacity Planning Becomes Multi-Dimensional

Traditional capacity planning may concentrate on ports, links, racks, and local growth. Scale-across AI adds geography, facility limits, and long-distance connectivity to that calculation.

A practical planning process should examine:

  1. Compute placement: Identify where additional accelerators can physically and electrically be accommodated.

  2. Local network capacity: Determine whether each site can support the planned compute density.

  3. Inter-site capacity: Treat long-distance connectivity as part of the overall fabric rather than as a separate transport layer.

  4. Dependency and failure analysis: Document which workloads and infrastructure components rely on each path or location.

  5. Operational headroom: Account for maintenance, configuration changes, and expansion instead of planning only for the initial deployment.

The key lesson is that compute, facility, and network planning can no longer happen independently. A site with available floor space is not useful if its power or network connectivity cannot support the intended role.

Operational Visibility Must Cross Site Boundaries

Distributed infrastructure also challenges tools and processes built around individual data centers. An isolated device view cannot explain the full impact of a change within a scale-across fabric.

Operations teams need visibility into:

  • Which network assets participate in the distributed environment

  • How devices and sites are connected

  • Whether configurations have drifted between locations

  • What changed before an incident or performance issue

  • Which assets have known vulnerabilities or are approaching end of life

  • Whether infrastructure configurations continue to satisfy internal and regulatory requirements

The objective is not merely to collect more device data. It is to preserve the relationships between assets, configurations, topology, lifecycle status, and change history across the entire multi-site fabric.

Where ConnectMyAssets Fits

Scale-across networking is a vendor innovation and architectural direction. ConnectMyAssets provides the vendor-agnostic, on-prem management layer needed to keep the underlying multi-vendor network estate understandable and controlled.

Relevant modules include:

  • Dynamic CMDB and Topology: Maintain an inventory of participating network assets and visualize their relationships across locations.

  • Backup & History: Version device configurations, investigate what changed, and use one-click rollback when a supported configuration must be restored.

  • Automation & ZTP: Apply repeatable network changes across distributed infrastructure while reducing manual inconsistency.

  • Compliance Engine: Assess configurations against frameworks including NIS2, ISO 27001, PCI, CISA, and NIST.

  • Per-asset CVE Tracking and End-of-Life Tracking: Connect security and lifecycle information to the devices supporting the AI fabric.

  • Credential Vault and SSH Bastion: Centralize controlled administrative access without sending infrastructure credentials to a cloud management service.

  • Local AI Insights: Analyze infrastructure information on premises, supporting operational review while keeping management data local.

ConnectMyAssets does not schedule AI workloads or replace an AI fabric’s own control mechanisms. Its role is to manage the network infrastructure beneath that fabric—providing inventory, topology, configuration history, compliance, lifecycle context, and controlled automation across vendors and sites.

Preparing for the Scale-Across Era

Before extending an AI environment across geographies, network teams should establish a reliable baseline:

  • Inventory every network asset involved in local and inter-site connectivity.

  • Map dependencies between sites, paths, and device configurations.

  • Back up configurations and retain an auditable change history.

  • Identify configuration drift before it becomes a cross-site problem.

  • Include vulnerability and end-of-life status in capacity and expansion decisions.

  • Define repeatable automation for changes that must remain consistent across locations.

Scale-across fabrics promise a practical response to the physical limits of centralized AI infrastructure. They also make disciplined network management more important: when compute spans geographies, the network becomes the connective foundation of the AI system itself.


Source: Arista Networks

Share this articleLinkedIn ↗Email ↗

Keep exploring.

All articles
Network management

Inside Highmark Stadium’s Converged Network: Lessons for Resilient Venue IT

The Buffalo Bills replaced the fragmented infrastructure of their former stadium with a converged Cisco network supporting Wi-Fi, broadcast, digital signage, communications, and location analytics. Highmark Stadium offers IT leaders a practical case study in service integration, operational ownership, wireless design, and the controls required when many critical functions share one foundation.

Read article
Network management

Autonomous AI in Network Operations: Trust Requires Evidence

A Cisco and Omdia survey suggests that network teams are increasingly willing to let AI take production actions—but almost never without guardrails. Before expanding autonomy, organizations need defined scopes, approval controls, traceable changes, and reliable rollback evidence across every managed asset.

Read article