Engineering
Site Reliability Engineer (SRE) Job Description
About the role
We are hiring a Site Reliability Engineer (SRE) to keep our production platform fast, available, and resilient as it scales. You will own service level objectives, build the automation and observability that catch problems before customers do, and lead incident response when things break. This role is a fit for someone who codes their way out of toil and treats reliability as a product feature, not an afterthought.
Key responsibilities
- Define and defend SLOs and error budgets for critical services
- Automate away toil and reduce manual operational work
- Lead incident response and drive blameless postmortems
Responsibilities
- Set service level objectives and track error budgets across critical services
- Build monitoring, alerting, and dashboards in tools like Prometheus, Grafana, and Datadog
- Automate deployments, scaling, and recovery using Terraform, Ansible, and CI/CD pipelines
- Lead on-call rotations and coordinate incident response during outages
- Write blameless postmortems and track corrective actions to closure
- Run capacity planning and load testing ahead of traffic spikes
- Harden systems for high availability with redundancy and graceful degradation
- Tune Kubernetes clusters, autoscaling, and resource limits for cost and performance
- Instrument services with distributed tracing and structured logging
- Partner with product engineers to build reliability into new features from the start
Requirements
- Bachelor's degree in computer science, engineering, or equivalent hands-on experience
- 3 or more years in SRE, DevOps, or production systems engineering
- Strong coding ability in a language such as Go, Python, or Rust
- Deep experience operating Linux systems and container orchestration at scale
- A track record of reducing incidents and improving uptime with measurable results
Make this JD your own
Generate a tailored Site Reliability Engineer (SRE) description with your company details, tone and must-have skills in seconds.
Keep your hiring moving
Hiring a Site Reliability Engineer (SRE)?
Send one link. Candidates record answers on their own time and AI ranks your shortlist, no scheduling, no back-and-forth.