AI-driven networking and infrastructure design at APNIC 62

By on 30 Sep 2026

Category: Tech matters

Tags: , , ,

Blog home

Sid Mathur discusses authoring IP geofeeds using AI tools and MCP, during Technical Session 2 at APNIC 62.

As networks evolve to support AI workloads, operators are rethinking both physical infrastructure and how they manage increasingly complex environments. At APNIC 62, during Technical Session 2 — AI in Networks, speakers explored how advances in data centre architecture, Machine Learning, agentic AI, and operational automation are reshaping network design and operations.

Topics ranged from next-generation optical fabrics and autonomous network management to predictive fault detection and AI-assisted geofeed authoring, highlighting both the opportunities and challenges of building and operating networks at scale in the AI era. This post covers the individual presentations from the session.

From fat tree to jellyfish: Rewiring India’s networks for the AI era

Lalit Singh Chowdhary, Chief Technology & Innovation Officer, Lightstorm Telecom Connectivity Pvt Ltd

Lalit Singh Chowdhary examined the often-overlooked infrastructure that underpins modern data centre design. In 2019, around 200 million Ethernet ports were sold for use inside switching fabrics. Today, that figure is closer to one billion, with 40% operating at 400Gbps or faster. There are now more active network ports than babies being born.

The concept of the ‘fat tree’ originated in 1985 through the work of Charles Leiserson at MIT. As traffic moves up the switching hierarchy, available bandwidth increases to prevent bottlenecks. This design underpins the traditional switch-edge-aggregation-core architecture, also known as the spine-and-leaf model.

At the core of the network, oversubscription is not acceptable. These layers are designed with a 1:1 ratio. At the edge, however, some oversubscription can be tolerated.

Ultimately, data centre architecture depends on three factors: The number of equal-cost paths available, the time required to traverse those paths, and the level of oversubscription that can be accepted while maintaining five nines availability. Historically, networks were designed for front-end workloads, reflecting the capacity and metro-area latency requirements of the time.

Lightstorm uses a three-spine, fully meshed interconnect architecture, with growth occurring primarily at the leaf layer. Any connected data centre can reach another within two hops. Interconnects are configured as wavelengths, and traffic is distributed across the three-spine design to achieve five nines availability, equivalent to around 26 seconds of downtime per month. If a path fails, traffic can be redistributed without requiring large-scale workload migration between data centres.

Cost remains a challenge. Operators must increasingly move from 10Gbps and 40Gbps line cards to 100Gbps and 400Gbps platforms. This shift can be difficult for customers accustomed to older rack and interconnect cost models.

At the same time, growing application complexity, sharding requirements, and fan-out patterns make workload placement decisions increasingly difficult.

Lalit then introduced an alternative approach. The ‘jellyfish’ architecture, developed by researchers at the University of Illinois Urbana-Champaign and HP Labs, replaces structured tiers with a conceptually random topology that has no visible core.

The design uses colourless, directionless, contentionless reconfigurable optical add-drop multiplexer (CDC-ROADM) technology and can reduce path costs through wavelength switching. The wavelength-switching plane is fully transparent, allowing any wavelength to travel in any direction without contention. When a required wavelength is already in use, the system can dynamically recolour traffic.

This architectural shift is driven by expectations of a tenfold increase in capacity demand. Future growth is expected to be constrained increasingly by power availability, requiring workloads to move to locations where power is available and network capacity to follow. The architecture maintains five nines availability through a fully integrated optical interconnect operating under a single management framework.

View the slides.

AI in network operations

Akhilesh Thakur, HPE Networking

Akhilesh Thakur outlined a vision for autonomous networks with minimal manual intervention, reducing the operational burden associated with reactive troubleshooting.

Previous AI initiatives in network management often struggled because insufficient data was available. Today, network operators have access to far richer datasets spanning network elements, control and data planes, infrastructure platforms, business systems, and external information sources.

Common AI use cases include:

  • Task automation
  • Dynamic network reconfiguration
  • Event correlation
  • Event management
  • Digital twin modelling

Akhilesh discussed agentic AI, which breaks complex problems into smaller actions while maintaining context through a retrieval-augmented generation (RAG) knowledge base.

The Model Context Protocol (MCP) provides a standard interface between AI agents and operational tools. Rather than integrating directly with APIs, databases, or NetConf implementations, MCP acts as an intermediary layer. This gives agentic systems a consistent operational view across platforms and vendors while reducing integration complexity.

Akhilesh demonstrated a Junos MCP implementation. Following HPE’s acquisition of Juniper Networks, Junos routing and switching platforms can be integrated into large-scale AI and data centre environments. Large language models such as Claude can interact with those environments through Junos MCP.

MCP also supports feature introspection, allowing systems to identify available capabilities dynamically. This makes it possible to map high-level requests into executable workflows using natural language.

One example showed a network audit generated through a Claude conversation over MCP, including recommendations on reporting structure and analysis. Akhilesh emphasized the importance of strong negative prompts to reduce hallucinations.

A second demonstration integrated MCP with external sources such as Cloudflare Radar and RIPEstat to support the detection and investigation of BGP route hijacking.

He also described higher-order MCP models that combine information from multiple sources into a single analytical workflow. For example, latency data from hosts, routers, and switches can be analysed together to identify root causes. The agent constructs a unified model using tabular datasets, joins, and integrated testing, including ping and round-trip time measurements.

Another demonstration showed agentic AI performing network reconfiguration with human oversight. Fully autonomous operation remains a future objective, and current deployments maintain a human-in-the-loop approach.

For Method of Procedure (MOP) generation, the LLM receives the change objective and refines it using information from the RAG knowledge base. Using a site-wide software upgrade as an example, Akhilesh noted that combining inference and RAG within a multi-agent framework produces MOPs that align more closely with intended outcomes while reducing unnecessary actions.

Digital twin environments further support ‘what if’ analysis while minimizing risk to production systems.

However, these approaches introduce new risks, including hallucinations, credential exposure, expanded operational blast radius, and model drift. Effective governance requires current RAG data, strong auditing practices, and careful oversight.

Discussion from the audience focused on industry caution regarding automation, the combined cost of human oversight and AI token consumption, testing methodologies, and the performance of open-source models.

View the slides.

Live ML solution for network AI Ops of broadband services

Dr Girish Saraph, Vegayan Systems Pvt. Ltd.

Dr Girish Saraph discussed the widening gap between network growth and the resources available to small and medium-sized enterprises and Network Operations Centres (NOCs). Increasing expectations for responsiveness require operators to prioritize issues more effectively and identify potential failures before they occur.

The example network contained 10 million customer endpoints across optical network terminal (ONT) and gigabit passive optical network (GPON) infrastructure. It generated 10 million alarms, approximately 200 active issues each day, and required four to six hours on average from diagnosis to remediation. The scale exceeded the capacity of available NOC staff.

Prioritization is critical because missing a major outage can trigger substantial service-level agreement penalties. Predictive and proactive insights can therefore provide significant operational value.

The ML Ops platform acts as a network expert, drawing data from either a data lake or directly from operational systems. Historical outage and fault records spanning several years are used to train multiple specialist models, each focused on a specific problem domain. These models process live network data and generate predictive insights and prioritization recommendations.

Applied to ONT and PON environments, the platform identifies ports with a history of recurring faults. In the example presented, 7.5% of customer access ports fell into this category. The system achieved a reported 74% hit rate when predicting reviewed ports, improving customer-facing reliability. Early identification allows operators to remediate issues before customers are affected, reducing churn.

The same technique was applied to GPON aggregation nodes. In the example network, 5.8% of nodes were classified as high risk, with reported predictive accuracy of 65%. Operators gained a three- to four-day prediction window and could better distinguish fibre cuts from hardware-related issues such as overheating or voltage instability. The approach reportedly improved service performance by 20%.

Given the volume of alarms generated in modern networks, the platform consolidates alarms into events, assigns severity levels, and ranks incidents according to customer impact. This allows NOC teams to focus on the highest-priority issues.

Dr Saraph emphasized that no single Machine Learning algorithm should be assumed to be optimal. Operators should test multiple approaches and select the model that performs best for their specific environment.

The presentation demonstrated how predictive analytics can improve reliability, reduce operational costs, and strengthen customer retention in ONT and GPON networks.

View the slides.

Authoring IP geofeeds using AI tools and MCP

Sid Mathur, Fastah Inc.

Sid Mathur’s presentation covered geofeed publication best practices and the use of MCP-enabled AI tools to improve geofeed quality.

He began with the example of a customer whose IP address appeared to be located in a different economy because of incorrect geolocation information. In many cases, content providers direct these complaints back to the customer’s ISP.

Internet Service Providers publish geolocation information for IP address blocks using the format defined in RFC 8805. The format resembles a CSV file and emerged from industry and regulatory requirements to identify the location of IP address allocations.

The model is straightforward. Operators associate an IP prefix with an ISO 3166 economy code, a region identifier, and a city. All fields are optional, allowing publication at economy, regional, or city level.

Operators can also publish ‘NONE’, indicating that addresses should not be geolocated. This is particularly useful for infrastructure addresses and anycast deployments, where assigning a specific location may be misleading.

RFC 9632 defines discovery mechanisms within Regional Internet Registry frameworks, enabling geofeed consumers to locate and interpret published data. RFC 9877, currently under development, proposes representing the same information using a JSON-based Registration Data Access Protocol (RDAP) model.

Sid noted that geofeed management is an ongoing process rather than a one-time activity. Smaller ISPs may update data infrequently, while larger networks may require continual changes as allocations evolve.

Publishing geofeeds through an RIR-linked URL places responsibility for data accuracy on the address holder. Platforms such as GitHub can host the data, but external parties do not control the content. Because accurate geolocation information has privacy implications, city-level accuracy is typically sufficient and should be viewed as approximate rather than precise.

Sid then explored the use of LLMs within geofeed workflows. When provided with relevant RFC guidance and IP address management (IPAM) information, an LLM can generate compliant geofeeds and perform audits.

He has developed an AI skill focused on geofeed management. Demonstrations showed the tool identifying common problems in public geofeeds, including missing region information for ambiguously named cities such as Frankfurt. It can also detect excessive use of ‘do not geolocate’ entries and identify cases where numeric values have been used instead of ISO 3166 codes. These checks function much like a software linter.

An MCP server can provide access to specialized geographic information, helping distinguish locations that share the same name. Mapping tools and scale-independent visualizations support this process.

Sid concluded by recommending the GitHub-hosted CSV publication model and demonstrated how his semantic validation tool can be integrated directly into operational workflows.

View the slides.

A common theme across these presentations was the growing need for intelligence, flexibility, and automation throughout the network stack. Whether through new data centre topologies designed for massive AI-driven workloads, AI agents capable of assisting with network operations, Machine Learning systems that predict faults before they occur, or tools that improve the quality of operational data, the goal is increasingly the same: Enabling operators to manage more complexity with greater confidence.

While challenges around cost, governance, accuracy, and trust remain, the technologies presented at APNIC 62 demonstrated how AI and advanced automation are becoming practical tools for improving network performance, resilience, and operational efficiency.

Watch the session in full:


The views expressed by the authors of this blog are their own and do not necessarily reflect the views of APNIC. Please note a Code of Conduct applies to this blog.

Leave a Reply

Your email address will not be published. Required fields are marked *

Top