> Closing the observability gap for the AI-ready enterprise
[DATE: 23/09/2026 15:00]
[LANGUAGE: EN]
The modern enterprise is a digital enterprise. From the back office to the factory floor, connected systems and digital services form the operational backbone on which all else depends. So when disruption hits, it can have a huge financial, reputational, productivity, and even compliance impact. This has raised observability to a board-level issue. "For a public company, a material cyber incident is a disclosure obligation. You're on a four-business-day clock from the moment you determine its material," explains NETSCOUT director of enterprise strategy, Jack Callahan. "So when you have a disruption, whether that's a cyber-attack, a DDoS attack, or someone pushing a bad update to the network, the first executive problem is the same: figuring out whether it’s material." With each technical team pointing fingers at each other, observability becomes the single source of truth that organizations need to identify root cause, accelerate resolution, and improve reliability. Yet in many enterprises, it’s not having the desired impact. The long-established data foundation of metrics, events, logs, and traces (MELT) can’t by itself keep pace with the complexity and scale of today’s digital infrastructure. Organizations have defaulted to gathering more data, increasing sampling, and extending retention. But they’re not getting better insight. “Executives who would expect to have a lot of data in front of them with which to make a decision don't always find that that data is as conclusive as they'd want it to be,” Callahan continues. “And therefore, they’re trusting their gut more than they’d expect, given how much they’re spending.” The costs of this observability debt are building. One study by NETSCOUT reveals that 81 percent of organizations believe insufficient data increases incident resolution time. Over two-fifths (42 percent) estimate downtime at $500,000 to$999,000 per hour. These costs are unsustainable, both economically and otherwise. To harness the power of autonomous AI in operations, organizations need a data foundation they can trust implicitly. This demands a fresh approach; economically viable and grounded in observability data that’s consistent, comprehensive, enriched, and real time. And delivered in a way that complements rather than replaces existing observability investments — extending the value of the platforms already embedded in the enterprise stack. Where visibility fails MELT data is still essential to observability. But it wasn’t designed for today’s complex, distributed and dynamic operations. Metrics explain that something has changed over time. Events surface when something changed. Logs tell teams that something happened at a specific time. But they don’t provide the context that explains what actually happened on a network and why. Traces come closest, as distributed tracing is built to follow a request across services. But a trace only shows what has been instrumented, which leaves it blind at un-instrumented components, third-party dependencies, and the infrastructure in between. And those are exactly where things tend to break down, meaning the context of what actually happened and why isn’t captured. Context essentially means being able to reconstruct a single, complete and ordered chain of events across different systems — including what kick-started an event, how it propagated, and what happened at each step. This is where MELT-only observability techniques often fail. Timestamps can be inconsistent across different systems. Identifiers might not be preserved across architectural boundaries. Sampling and aggregation remove vital detail needed for reconstruction. And data may be stored across different tools with incompatible schemas. Research reveals that 96 percent of organizations use metrics and logs, yet 82 percent report visibility gaps, and nearly all (96 percent) lack sufficient data to determine root cause during incidents. They tend to lose visibility where systems meet, such as between on-premises and cloud (58 percent), the edge (51 percent), or in service-to-service interactions (39 percent). AI sharpens the challenge These issues become more serious in an AI context. Organizations are already embracing AI-driven operations to improve efficiency, decision making and customer experiences. But when systems start operating autonomously, making decisions and taking action at machine speed, they need forensic-grade data with high-fidelity context to produce reliable outcomes. That means continuous, unsampled records that preserve system interactions across environments. Higher levels of autonomy demand higher levels of confidence in network data. But telemetry can lose fidelity through sampling and abstraction — common techniques used in MELT to manage high data volumes. The resulting incomplete and fragmented data can lead to false correlation, ambiguity over root cause, inconsistent outputs, and overconfidence in partial signals. “An agent is not going to apply human intelligence to troubleshoot an issue. It's going to make a decision based on the data it has,” says Callahan. “So if you are feeding it partial, or periodic, or sampled data, you're at risk of scaling that uncertainty really quickly.” It’s a challenge that many organizations are just waking up to. According to NETSCOUT, only 41 percent describe AI-assisted insights as “very or extremely consistent.” A similar share (38 percent) admits to lacking forensic-grade data to validate automated actions. Some 29 percent say they don’t have real-time visibility across environments, and 28 percent don’t fully trust automation output. Closing the observability gap A better approach would be to build observability around MELT data enriched to provide the context that IT teams need, but without the bloat that adds unsustainable extra cost. This starts with packet data: the authoritative record of what actually traversed the network. It provides visibility into the transactions, dependencies and interactions (human and machine-based) across the IT ecosystem. Using deep packet inspection (DPI) techniques, this visibility can be distilled into metadata that, added to MELT, produces what NETSCOUT calls “MELT+”. “Digital services become observable through the exchanges among their components. NETSCOUT Smart Data transforms those observed interactions into transaction-level evidence: whether communication succeeded, how the transaction performed, where delay or failure appeared, which services were affected and, when identity context is available, which users experienced the impact. That gives operations teams and AI systems a more complete and trustworthy basis for understanding what actually happened,” explains NETSCOUT field marketing manager , Steve Horneman. “Most telemetry describes the state of individual components. NETSCOUT observes the interactions among those components and creates meaning from them as the activity occurs. By extracting context early, from independently observed traffic rather than relying only on what individual systems report, we give operations platforms and AI a more consistent account of how a digital service actually behaved. That is the difference between collecting more telemetry and creating evidence that can support a confident decision.” One case illustrates the advantage of this approach. A product manufacturer found that wireless connectivity issues were causing automated guided vehicles (AGVs) to fail in its global facilities, costing the company $500,000 per hour in lost productivity. Outages were occurring roughly every three weeks. Existing robotics telemetry failed to find the root cause. But once NETSCOUT was pulled in, the source of the issue was pinpointed, and a proactive monitoring model adopted which detects AGV failures within seconds. Troubleshooting fell from hours to minutes, saving the company tens of millions of dollars annually. The benefits of MELT+ expand beyond outages and operational incidents to cybersecurity, Horneman continues. “The strategic value extends beyond observability. The same independently observed interaction evidence can support operational assurance at the enterprise perimeter, expose service-to-service behavior and potential lateral movement internally, and give operations, security, and AI systems a common evidentiary foundation. Instead of each team interpreting a different version of events, they can reason from the same observed reality,” he says. NETSCOUT calculates that organizations treating network traffic data as authoritative are nearly three times more likely to report that visibility gaps occur infrequently (50 percent vs.18 percent). It is this level of insight into what’s happening on the network that makes the same packet-derived intelligence valuable to forensic analysis teams. “Once an attacker has privilege on a host, the telemetry that host generates about itself is within reach,,” says Callahan. “Sophisticated attackers hide lateral movement exactly that way. What they can't do is go back and change the packets that already crossed the network. That's a higher level of veracity, and a more complete view.” When metadata is Smart Data NETSCOUT’s approach uses DPI to observe live, unsampled packets directly from the network and then convert it into high-fidelity metadata using Adaptive Service Intelligence (ASI). It’s designed to tackle the main challenges of traditional MELT: scale, efficiency, cost, and data richness. NETSCOUT observes traffic from strategic points in the network rather than monitoring each application or server, reducing telemetry volume, ingestion cost, and complexity. It analyzes and distills packet data into Smart Data, metadata generated at the point of capture, which reduces the volume that needs to be moved, stored or retained downstream. What customers get is an approach that is complementary to MELT but which is economically more sustainable, produces more complete, network-derived data, and which feeds into existing observability platforms to further reduce TCO. It also delivers what analyst firm Futurum describes as the critical foundation for autonomous AI operations. Data that captures verifiable network behavior and observed interactions rather than abstractions. Data that ensures comprehensive visibility regardless of whether individual applications have been instrumented, and a complete view without sampling gaps. And which is consistent across observability, security, and operations teams, while demonstrating sequence and causality across service boundaries. “MELT alone is not going to be a sufficient data foundation to run AIOps on,” says Callahan. “We're able to generate data with more of the context you need earlier in the process, and therefore richer data flows into your platforms.” Just getting started Despite the obvious benefits of MELT+ approaches, NETSCOUT data reveals that only 11 percent of organizations treat full-fidelity network data as authoritative. For CIOs keen to change that statistic, the first step is to evaluate their current observability data by five key criteria, as shared by NETSCOUT COO, Sanjay Munshi. It should be comprehensive; covering any cloud, service, app, network or vendor. It should be curated; with purpose-built feeds optimized for storage and cost. It must be credible in offering a verifiable chain of interactions showing how services, apps and users behave in context. It must be consistent across use cases. And it must provide continuous real-time insight into data in motion. “If you’ve been optimizing to reduce your MELT cost, what you’ve been doing is also reducing the context that your application teams and agents have. But you no longer have to sacrifice one in order to gain the other,” Callahan concludes. “If you’re worried about telemetry costs. If you're worried about having the data you need to make decisions in the moment or for compliance reporting. If you're trying to figure out how to move your AI pilots into production: we can strengthen what you are already doing in the platforms you use every day.” Sponsored by NETSCOUT