Detection is only half the process. Swift remediation is what keeps operations smooth. Yet for many site reliability engineers (SREs), on-call shifts are spent trapped in a loop of repetitive firefighting. Pairing Splunk Observability Cloud with Red Hat Ansible Automation Platform helps bridge the gap between discovery and response, transforming manual triage into safer, automated self-healing that reduces mean time to resolution (MTTR).

The architecture

AIOps uses operational alerts to automate routine operational tasks. For example, when Splunk Observability Cloud detects a memory leak, a certificate nearing expiration, or an application throwing errors, Ansible Automation Platform can execute the same remediation steps an engineer would, such as restarting a service, scaling a deployment, or rotating a credential, in seconds instead of hours.

Four components form the pipeline:

  • Splunk OpenTelemetry Collector: Gathers metrics, traces, and logs from your hosts, containers, and Kubernetes clusters, routing them directly to Splunk Observability Cloud.
  • Splunk Observability Cloud: Aggregates metrics, traces, and logs, applying AI-driven detectors to catch performance degradation before users do.
  • Webhook: Sends Splunk alerts as structured event payloads to Ansible Automation Platform either as native webhooks or through the Splunk add-on for Event-Driven Ansible.
  • Ansible Automation Platform: Receives events through the Event-Driven Ansible feature, matches them against a rulebook, and executes remediation playbooks with full role-based access control (RBAC) and audit trails.

Deploy and manage collectors at scale with Splunk OTel and Ansible Automation Platform

Managing 1 collector on a test box is easy. Deploying hundreds across a hybrid fleet with consistent pipelines and configs is where manual work falls apart.

The Splunk OpenTelemetry (OTel) Collector simplifies this process by acting as your universal telemetry agent, gathering data from hosts, containers, and applications, and routing it cleanly to Splunk Observability Cloud. The cisco.splunk_otel_collector Ansible Content Collection (available on Ansible automation hub) handles the full lifecycle:

  • Deploy across Linux, Windows, and container environments in a single run.
  • Configure receivers, processors, and exporters consistently fleet-wide.
  • Update with controlled batching and canary strategies.
  • Validate collectors are healthy and reporting post-deployment.

Ansible Automation Platform is the governed execution layer that uses RBAC for sensitive configurations, scheduled rollouts, credential vaulting, and simplified surveys. This allows operators to customize deployments without editing playbook code, turning a high-stress, manual rollout into a predictable, repeatable, and trusted operation.

What Splunk Observability Cloud monitors

With collectors deployed, Splunk Observability Cloud provides correlated visibility across the full environment:

  • Infrastructure Monitoring: Host and container metrics from the OTel collectors (CPU, memory, disk, and network) with real-time dashboards and alerts.
  • Application Performance Monitoring (APM): Distributed traces across services, mapping dependencies, and surfacing latency bottlenecks.
  • Log Observer Connect: Logs correlated with metrics and traces so operators get full context without switching tools.
  • Real User Monitoring (RUM): Frontend performance from real user sessions including page loads, JavaScript errors, and experience degradation caused by backend changes.
  • Synthetic Monitoring: Proactive availability and performance tests from external vantage points, catching outages before users report them.
  • On-Call and Incident Response: Alert routing and incident workflows.

These observability signals aren't just for dashboards, they serve as the triggers for automated remediation downstream.

From alert to remediation: How the loop works

Here's what this looks like in practice with a memory capacity alert example.

  1. Deploy and detect. Ansible Automation Platform deploys the Splunk OTel Collector. Memory on your application servers crosses 90%. Splunk Observability Cloud sees garbage collection pause times climbing and response latency rising, firing an alert.
  2. Event forwarding. The webhook sends the alert payload (hostnames, memory values, severity) directly to Ansible Automation Platform as a structured event.
  3. Event received. The Event-Driven Ansible feature included in Ansible Automation Platform picks up the event.
  4. Rule matching. Ansible Automation Platform evaluates the event against your existing, pre-approved Ansible Rulebooks.
  5. ITSM sync (optional). Before running the fix, Ansible Automation Platform can create and update IT service management (ITSM) tickets via ITSM integration using the platform API. Once the automation runs, the ticket is updated and closed if successful.
  6. Rule evaluation. Once approved (or automatically validated by policy), the rulebook passes alert context, such as hostnames and breached thresholds, directly into the remediation playbook.
  7. Execute with approval. The Ansible Playbook or job template is called. This step can include a human-in-the-loop (HITL) where an engineer's approval is required.
  8. Remediate. The Ansible Playbook or job template runs and safely restarts the application process with a rolling strategy, 1 host at a time, validating health checks between each restart. If the remediation fails (for example, if memory use spikes again within minutes), a different rule is activated. This immediately escalates the incident to the on-call engineer using the ITSM platform, providing full context on what was tried, the automated results, and the current state of the environment.
Diagram illustrating a continuous AIOps architecture loop connecting Splunk Observability Cloud and Red Hat Ansible Automation Platform through instrumentation, observability, and automated remediation."

What about the playbook you don't have yet?

Not every outage has a pre-built playbook waiting. With AI-assisted authoring tools, whether an automation coding assistant (a feature within Ansible Automation Platform), the Model Context Protocol (MCP) server for Ansible Automation Platform, or a team's own large language model (LLM) tooling such as OpenAI or Claude, engineers can describe remediation steps with a natural language prompt and generate targeted Ansible Automation Platform tasks in seconds. This empowers additional teams to use automation and accelerates productivity.

Getting started

  1. Instrument: Access the cisco.splunk_otel_collector Ansible Content Collection available on Ansible automation hub as part of your Ansible Automation Platform subscription.
  2. Observe: Configure Splunk Observability Cloud detectors for your critical services. Wire Ansible Automation Platform's own telemetry into Splunk so automation health is also visible.
  3. Automate: Download the Splunk Add-on for Event-Driven Ansible on Splunkbase or create lightweight webhooks to automate your landscape. You can explore sample rulebooks to start automating your operational workflows.

Resources

Product trial

Red Hat Ansible Automation Platform | Product Trial

An agentless automation platform.

About the author

Stephen Fulmer is a Product Manager at Red Hat, leading Ansible content strategy. With a background in virtualization and IT operations, he works closely with customers, partners, and engineering teams to deliver trusted, scalable automation content for platforms like OpenShift, Windows, and public cloud. Stephen is passionate about enabling organizations to simplify complex workflows and accelerate their automation journeys with Red Hat Ansible Automation Platform.

UI_Icon-Red_Hat-Close-A-Black-RGB

Keep exploring

Browse by channel

automation icon

Automation

The latest on IT automation for tech, teams, and environments

AI icon

Artificial intelligence

Updates on the platforms that free customers to run AI workloads anywhere

open hybrid cloud icon

Open hybrid cloud

Explore how we build a more flexible future with hybrid cloud

security icon

Security

The latest on how we reduce risks across environments and technologies

edge icon

Edge computing

Updates on the platforms that simplify operations at the edge

Infrastructure icon

Infrastructure

The latest on the world’s leading enterprise Linux platform

application development icon

Applications

Inside our solutions to the toughest application challenges

Virtualization icon

Virtualization

The future of enterprise virtualization for your workloads on-premise or across clouds