* [Topics](/en/topics "Topics")
* [DevOps](/en/topics/devops "DevOps")
* What is SRE (site reliability engineering)?
What is SRE (site reliability engineering)?
===========================================
Published  May 4, 2020•*4*-minute read
Copy URL
Jump to section
---------------
OverviewWhat does a SRE do?DevOps vs. SREPlatform engineer vs. SRETechnology to support SRE
Overview
--------
Site reliability engineering (SRE) is a software engineering approach to IT operations. SRE teams use software as a tool to manage systems, solve problems, and [automate](/en/topics/automation) operations tasks.
SRE takes the tasks that have historically been done by operations teams, often manually, and instead gives them to engineers or operations teams who use software and automation to solve problems and manage production systems.
SRE is a valuable practice when creating scalable and highly reliable software systems. It helps  manage large systems through code, which is more scalable and sustainable for system administrators (sysadmins) managing thousands or hundreds of thousands of machines.
The concept of site reliability engineering comes from the Google engineering team and is credited to Ben Treynor Sloss.
SRE helps teams find a balance between releasing new features and ensuring reliabilty for users.
In this context, standardization and automation are 2 important components of the SRE model. Here, site reliability engineers seek to enhance and automate operations tasks.
In these ways, SRE helps improve system reliability today—and as it grows over time.
SRE supports teams that are moving their IT operations from a traditional approach to a [cloud-native](/en/topics/cloud-native-apps) approach.
[Learn about Red Hat's approach to SRE](/en/solutions/sre-approach "Topic | Red Hat’s approach to SRE")
What does a site reliability engineer do?
-----------------------------------------
A site reliability engineer is a unique role that requires either a background as a sysadmin, a software developer with additional operations experience, or someone in an IT operations role that also has software development skills.
SRE teams are responsible for how code is deployed, configured, and monitored, as well as the availability, latency, change management, emergency response, and capacity management of services in production.
SRE teams determine the launch of new features by using service-level agreements (SLAs) to define the required reliability of the system through service-level indicators (SLI) and service-level objectives (SLO).
An SLI measures specific aspects of provided service levels. Key SLIs include request latency, availability, error rate, and system throughput. An SLO is based on the target value or range for a specified service level based on the SLI.
An SLO for the required system reliability is then based on the downtime determined to be acceptable. This downtime level is referred to as an error budget—the maximum allowable threshold for errors and outages.
With SRE, 100% reliability is not expected—failure is planned for and expected.
Once established, the development team is able to "spend" the error budget when releasing a new feature. Using the SLO and error budget, the team then determines whether a product or service can launch based on the available error budget.
If a service is running within the error budget, then the development team can launch whenever it wants, but if the system currently has too many errors or goes down for longer than the error budget allows then no new launches can take place until the errors are within budget.
The development team conducts automated operations tests to demonstrate reliability.
Site reliability engineers split their time between operations tasks and project work. According to SRE best practices from Google, site reliability engineers can only spend a maximum of 50% of their time on operations—and they should be monitored to ensure they don’t go over.
The rest of the their time should be spent on development tasks like creating new features, scaling the system, and implementing automation.
Excess operational work and poorly performing services can be redirected back to the development team so that the site reliability engineer doesn't spend too much time on the operations of an application or service.
Automation is an important part of the site reliability engineer’s role. If they are repeatedly dealing with a problem, then they will likely automate a solution.
Maintaining the balance between operations and development work is a key component of SRE.
Red Hat resources
-----------------
[Keep reading](/en/resources "Keep reading")
DevOps vs. SRE
--------------
[DevOps](/en/topics/devops) is an approach to culture, automation, and platform design intended to deliver increased business value and responsiveness through rapid, high-quality service delivery. SRE can be considered an implementation of DevOps.
Like DevOps, SRE is about team culture and relationships. Both SRE and DevOps work to bridge the gap between development and operations teams to deliver services faster.
Faster application development life cycles, improved service quality and reliability, and reduced IT time per application developed are benefits that can be achieved by both DevOps and SRE practices.
However, SRE differs from DevOps because it relies on site reliability engineers within the development team who also have an operations background to remove communication and workflow problems.
The site reliability engineer role itself combines the skills of development teams and operations teams by requiring an overlap in responsibilities.
SRE can help DevOps teams whose developers are overwhelmed by operations tasks and need someone with more specialized operations skills.
When coding and building new features, DevOps focuses on moving through the development pipeline efficiently, while SRE focuses on balancing site reliability with creating new features.
Here, modern application platforms based on container technology, Kubernetes and [microservices](/en/topics/microservices/what-are-microservices) are critical to DevOps practices, helping deliver security and innovative software services.
[Learn how to implement DevOps with a Kubernetes platform](https://go.redhat.com/accelerate-devops-openshift-20180918?intcmp=701f2000001OMH6AAO)
Platform engineer vs SRE
------------------------
Both [platform engineering](/en/topics/platform-engineering/what-is-platform-engineering) and site reliability engineering are about creating and maintaining systems. The difference between the two concepts lies in the focus of each practice. An SRE places their focus on IT operations teams, helping them use software as a tool to manage systems, solve problems, and automate operations tasks.
Platform engineers focus on development teams, helping them create platforms for managing systems, solving problems, and automating development tasks.
Technology to support SRE
-------------------------
SRE relies on automating routine operational tasks and standardization across an [application’s lifecycle](/en/topics/devops/what-is-application-lifecycle-management-alm). Red Hat® Ansible® Automation Platform is a comprehensive, integrated platform that helps SRE teams automate for velocity, collaboration, and growth—offering security and support across the technical, operational, and financial functions of the enterprise.
Specifically, Ansible Automation Platform offers:
* Infrastructure orchestration on cloud and on-premise for instances, routing, load balancing, firewalls, and more.
* Infrastructure optimization, including right-size cloud resources and adding or removing resources like central processing unit (CPU) and random access memory (RAM) as needed.
* [Cloud operations](/en/topics/automation/what-is-cloudops), including application deployments with continuous integration and continuous delivery (CI/CD) pipelines, operating system patching, and maintenance.
* Business continuity, including moving and copying resources off cloud, creating and managing policies for backups, and managing disruptions and failures.
[Red Hat Ansible Automation Platform: A beginner's guide](/en/resources/ansible-automation-platform-beginners-guide-ebook)
SRE also relies on a foundation designed for a cloud-native development style. [Linux® containers](/en/topics/containers/whats-a-linux-container) support a unified environment for development, delivery, integration, and automation.
And [Kubernetes](/en/topics/containers/what-is-kubernetes) is the modern way to automate Linux container operations. Kubernetes helps teams more efficiently manage clusters running Linux containers across public, private, or hybrid clouds.
As an enterprise-ready Kubernetes platform that supports SRE initiatives, [Red Hat® OpenShift®](/en/technologies/cloud-computing/openshift) helps teams implement culture and process transformation that modernizes IT infrastructure and positions organizations to better serve their customers and achieve business goals.
[Try Red Hat OpenShift for free](/en/technologies/cloud-computing/openshift/try-it "product | red hat openshift | try it")
The official Red Hat blog
-------------------------
Get the latest information about our ecosystem of customers, partners, and communities.
[Keep reading](/en/blog "The official Red Hat blog")
All Red Hat product trials
--------------------------
Our no-cost product trials help you gain hands-on experience, prepare for a certification, or assess if a product is right for your organization.
[Keep reading](/en/products/trials "All Red Hat product trials")
Keep reading
------------
### What is blue green deployment?
Blue green deployment is an application release model that gradually transfers user traffic from a previous version of an app or microservice to a nearly identical new release—both of which are running in production.
[Read the article](/en/topics/devops/what-is-blue-green-deployment "article | what is blue green deployment")
### What is observability?
Observability refers to the ability to monitor, measure, and understand a system or application by examining its outputs, logs, and performance metrics.
[Read the article](/en/topics/devops/what-is-observability "article | what is observability")
### How to approach DevOps metrics
DevOps metrics track the effectiveness of DevOps practices, which relate to software development and IT operations.
[Read the article](/en/topics/devops/how-approach-devops-metrics "article | how to approach devops metrics")
DevOps resources
----------------
### Related content
* Case study
  [Powering O2’s next-generation 5G network with Red Hat OpenShift](/en/resources/o2-czech-republic-case-study)
* Overview
  [Manufacturers combine modernization and continuity with Red Hat](/en/resources/anon-manufacturing-overview)
* Blog post
  [Strategic momentum: The new era of Red Hat and HPE Juniper network automation](/en/blog/strategic-momentum-new-era-red-hat-and-hpe-juniper-network-automation)
* Datasheet
  [Automate Google Cloud with Red Hat Ansible Automation Platform](/en/resources/automate-google-cloud-with-ansible-datasheet)
### Related articles
* [What is security automation?](/en/topics/automation/what-is-security-automation)
* [Why choose Red Hat for automation?](/en/topics/automation/why-choose-red-hat-for-automation)
* [What is an Ansible Playbook?](/en/topics/automation/what-is-an-ansible-playbook)
* [What is SOAR?](/en/topics/security/what-is-soar)
* [Learning Ansible basics](/en/topics/automation/learning-ansible-tutorial)
* [What is blue green deployment?](/en/topics/devops/what-is-blue-green-deployment)
* [How to build an IT automation strategy](/en/topics/automation/build-an-automation-strategy)
* [What is observability?](/en/topics/devops/what-is-observability)
* [Ansible vs. Puppet: What you need to know](/en/topics/automation/ansible-vs-puppet)
* [Ansible vs. Salt: What you need to know](/en/topics/automation/ansible-vs-salt)
* [Ansible vs. Chef: What you need to know](/en/topics/automation/ansible-vs-chef)
* [Ansible vs. Terraform](/en/topics/automation/ansible-vs-terraform)
* [What is IT service management (ITSM)?](/en/topics/automation/what-is-it-service-management-itsm)
* [How to approach DevOps metrics](/en/topics/devops/how-approach-devops-metrics)
* [Automating Microsoft Windows with Red Hat Ansible Automation Platform](/en/technologies/management/ansible/automate-microsoft-windows-with-ansible)
* [What is DevOps automation?](/en/topics/automation/what-is-devops-automation)
* [What is Infrastructure as Code (IaC)?](/en/topics/automation/what-is-infrastructure-as-code-iac)
* [What is application lifecycle management (ALM)?](/en/topics/devops/what-is-application-lifecycle-management-alm)
* [What is CI/CD?](/en/topics/devops/what-is-ci-cd)
* [Ansible vs. Kubernetes: how they work together](/en/topics/automation/Ansible-vs-Kubernetes)
* [What is a configuration management database (CMDB)?](/en/topics/automation/what-is-a-configuration-management-database-cmdb)
* [What is cloud migration? And how can automation help?](/en/topics/automation/what-is-cloud-migration)
* [What is DevOps?](/en/topics/devops/what-is-devops)
* [What is GitOps?](/en/topics/devops/what-is-gitops)
* [What is a software-defined data center (SDDC)?](/en/topics/automation/what-is-a-sddc)
* [What is a CI/CD pipeline?](/en/topics/devops/what-cicd-pipeline)
* [What is IT automation?](/en/topics/automation/what-is-it-automation)
* [Why choose Red Hat Ansible Automation Platform as your AI foundation?](/en/topics/automation/automation-and-ai)
* [Platform engineering vs. DevOps](/en/topics/platform-engineering/platform-engineering-vs-devops)
* [What is access control?](/en/topics/security/what-is-access-control)
* [What is virtual infrastructure management? And how can automation help?](/en/topics/automation/virtual-infrastructure-management)
* [What is IT migration?](/en/topics/automation/what-is-it-migration)
* [How to automate migrations with Red Hat Ansible Automation Platform](/en/technologies/management/ansible/automate-migrations-with-red-hat-ansible-automation-platform)
* [Why use Red Hat Ansible Automation Platform with Red Hat OpenShift?](/en/technologies/cloud-computing/openshift/ansible-on-openshift)
* [What is CloudOps?](/en/topics/automation/what-is-cloudops)
* [Red Hat Satellite on Red Hat Enterprise Linux](/en/technologies/management/satellite/satellite-for-rhel)
* [What is multi-cloud GitOps?](/en/topics/devops/what-is-multicloud-gitops)
* [What is role-based access control (RBAC)?](/en/topics/security/what-is-role-based-access-control)
* [What is a GitOps workflow?](/en/topics/devops/what-is-gitops-workflow)
* [Which Red Hat Ansible Automation Platform deployment option is right for you?](/en/technologies/management/ansible/ansible-deployment-options)
* [What is an Ansible module—and how does it work?](/en/topics/automation/what-is-an-ansible-module)
* [What is Argo CD?](/en/topics/devops/what-is-argocd)
* [How to manage and automate applications at the edge](/en/topics/edge-computing/how-to-manage-automate-applications-edge)
* [How to build an automation Center of Excellence](/en/topics/automation/how-to-build-automation-center-of-excellence)
* [What is orchestration?](/en/topics/automation/what-is-orchestration)
* [How to adopt Automation as Code: Extending Infrastructure as Code into Policy as Code](/en/topics/automation/how-to-adopt-automation-as-code)
* [What is a webhook?](/en/topics/automation/what-is-a-webhook)
* [Red Hat Lightspeed data and application security](/en/topics/management/data-application-security)
* [What is an Ansible Role—and how is it used?](/en/topics/automation/what-is-an-ansible-role)
* [What is CI/CD security?](/en/topics/security/what-is-cicd-security)
* [What is data management?](/en/topics/data-services/what-is-data-management)
* [Gain security with Red Hat Ansible Automation Platform](/en/technologies/management/ansible/gain-security-with-red-hat-ansible-automation-platform)
* [What is NetOps?](/en/topics/automation/what-is-netops)
* [What is an Ansible Rulebook?](/en/topics/automation/what-is-an-ansible-rulebook)
* [What is configuration management](/en/topics/automation/what-is-configuration-management)
* [What is event-driven automation?](/en/topics/automation/what-is-event-driven-automation)
* [Zero-Touch Provisioning and telco automation with Red Hat](/en/topics/telecommunications/zero-touch-provisioning-and-telco-automation-at-red-hat)
* [What is an internal developer platform?](/en/topics/platform-engineering/what-is-an-internal-developer-platform)
* [What is infrastructure automation?](/en/topics/automation/what-is-infrastructure-automation)
* [Why choose Red Hat for a DevOps Platform?](/en/topics/devops/why-choose-red-hat-for-devops)
* [What is provisioning?](/en/topics/automation/what-is-provisioning)
* [What is YAML?](/en/topics/automation/what-is-yaml)
* [Understanding Ansible, Terraform, Puppet, Chef, and Salt](/en/topics/automation/understanding-ansible-vs-terraform-puppet-chef-and-salt)
* [What is compliance management?](/en/topics/management/what-is-compliance-management)
* [What is cloud orchestration?](/en/topics/automation/what-is-cloud-orchestration)
* [What is a configuration file?](/en/topics/linux/what-configuration-file)
* [Ansible vs. Red Hat Ansible Automation Platform](/en/technologies/management/ansible/ansible-vs-red-hat-ansible-automation-platform)
* [What is cloud automation?](/en/topics/automation/what-is-cloud-automation)
* [What is agile methodology?](/en/topics/devops/what-is-agile-methodology)
* [What is network automation?](/en/topics/automation/what-is-network-automation)
* [What are managed IT services?](/en/topics/cloud-computing/what-are-managed-it-services)
* [What is business process management?](/en/topics/automation/what-is-business-process-management)
* [What is patch management (and automation)?](/en/topics/management/what-patch-management-and-automation)
* [What is the Red Hat Ansible Automation Platform automation controller?](/en/technologies/management/ansible/automation-controller-product-feature)
* [Cloud-native CI/CD on Red Hat OpenShift](/en/technologies/cloud-computing/openshift/ci-cd)
* [What is business process automation?](/en/topics/automation/what-is-business-process-automation)
* [What is IT process automation?](/en/topics/automation/what-is-it-process-automation)
* [What is continuous delivery?](/en/topics/devops/what-is-continuous-delivery)
* [What is deployment automation?](/en/topics/automation/what-is-deployment-automation)
* [What is business optimization?](/en/topics/automation/business-optimization)
* [What is Kubernetes cluster management?](/en/topics/containers/what-is-kubernetes-cluster-management)
* [What is risk management?](/en/topics/management/what-is-risk-management)
* [What is network management?](/en/topics/management/what-is-network-management)
* [What is IT system life-cycle management?](/en/topics/management/it-system-life-cycle-management)
* [What is an SOE?](/en/topics/management/what-is-an-soe)
* [What is robotic process automation (RPA?)](/en/topics/automation/what-is-robotic-process-automation)
* [What is cloud management?](/en/topics/cloud-computing/what-is-cloud-management)
* [What's business automation?](/en/topics/automation/whats-business-automation)
[More about this topic](/en/topics/devops "More about this topic")