IT Brief New Zealand - Technology news for CIOs & IT decision-makers
New Zealand
SRE teams take on wider AI oversight in production

SRE teams take on wider AI oversight in production

Wed, 26th Aug 2026 (Today)
Mark Tarre
MARK TARRE News Chief

Dynatrace has released research on the role of Site Reliability Engineering and platform engineering teams in managing AI in production. The survey found these teams are taking on broader responsibility for AI reliability and oversight.

The study of 919 IT leaders found that rapid AI adoption is changing what large enterprises expect from Site Reliability Engineering, or SRE, and platform engineering teams. Companies are increasingly relying on those teams to manage the operational demands of deploying AI systems at scale.

Among SRE respondents, 67% said monitoring AI models is now their top use case. The research also found that 58% use AI to monitor model performance and accuracy, while half use AI-powered tools for automated incident response.

Executive support for SRE appears strong. The research found that 92% of organisations reported executive leadership backing for SRE initiatives. Among organisations using platform engineering, 89% had implemented an internal developer platform and 60% reported broad adoption across departments.

These figures point to a broader shift in how AI is managed once it moves beyond development and into day-to-day operations. The research found that 73% of SRE and platform engineering teams now collaborate and share responsibilities across reliability and platform work.

Operational strain

The findings suggest many organisations are struggling to adapt existing monitoring and management approaches to AI workloads. AI systems create new telemetry requirements and introduce different failure patterns, adding to the complexity of running them reliably.

That complexity is showing up across the operating environment. Nearly half of SRE respondents said too many data sources and metrics made it harder to define and manage service-level objectives, while 37% of platform engineers said integration with existing tools and systems was their biggest challenge.

Observability remains uneven across deployment workflows. Only 40% of platform engineers said observability had been embedded across all deployment stages, indicating that many organisations still lack end-to-end visibility as AI systems move into production.

The report also suggested that AI has not yet delivered equally across all expected outcomes. While respondents said AI was generally helping with reliability and developer productivity, it was proving less effective at cutting costs and reducing mean time to resolution.

Governance focus

Some of the clearest signals in the research relate to governance and control. Teams are placing visibility and human oversight ahead of broader automation, reflecting caution about how autonomous systems should operate in live environments.

Service-level objectives remain a common practice in that context. The report found that 89% of SREs use service-level objectives across at least some teams or systems, suggesting companies are trying to apply established reliability measures to new AI-driven workloads.

Platform engineering teams are also taking on a larger role in AI adoption. More than half, or 55%, said they prioritise enabling developers with AI-powered tools such as coding copilots and chatbots, linking changes in developer workflows with broader platform oversight.

Dynatrace linked the findings to a wider industry move toward combining observability, automation and AI operations in a more unified model. The company has already announced its intention to acquire Arize, an AI observability specialist, as it looks to close the gap between teams building AI models and teams operating them in production.

That gap is becoming harder for enterprises to manage as AI systems become part of mainstream infrastructure rather than stand-alone experiments. The survey argued that fragmented tools and disconnected datasets are limiting organisations as they try to govern reliability across applications, infrastructure and AI systems together.

A Gartner forecast cited in the research said 80% of enterprises will adopt SRE practices across their organisations by 2028, up from 30% in 2024. That trajectory suggests the reliability operating model is likely to become more central as companies expand AI use.

Steve Tack, Chief Product Officer at Dynatrace, commented on the shift in responsibilities facing technical operations teams.

"SRE and platform engineering laid the groundwork for modern digital reliability, but AI is rewriting the rules. Enterprises now need to move from managing systems to orchestrating them, connecting observability, automation, and agentic AI to operate at the speed these initiatives demand, turning insight into action at scale," said Steve Tack, Chief Product Officer at Dynatrace.

He also addressed the separation between AI development and live operations.

"This research also reflects why we recently announced our intent to acquire Arize. AI engineering teams have been evaluating in one set of tools while operations teams monitor in another, and that gap is no longer sustainable as AI moves deeper into enterprise production," said Tack.