SRE: market
Site Reliability Engineer (SRE) — premium role inside the DevOps direction, invented by Google in 2003. Focus: reliability + SLI / SLO / error budgets + incident response + automation to reduce toil. Programming-heavier than general DevOps (Go / Python for automation + custom tooling). Role family: SRE (mid — owns reliability of one service), Senior SRE (multi-service + on-call mastery + SLO architecture), Staff / Principal SRE (org-wide reliability strategy + production engineering culture leadership), SRE Tech Lead (team + reliability roadmap), Production Engineer (alternative title — Facebook / Meta term). Stack 2026: Linux+bash deep mastery (production debugging), Go (standard for SRE automation — Kubernetes + Prometheus + most SRE tooling in Go), Python (data analysis + scripting), Kubernetes mastery (production-scale), Prometheus+Grafana+Alertmanager+VictoriaMetrics+Mimir mastery (metrics deep), Loki+Tempo+OpenTelemetry (logs + traces), Datadog/New Relic/Dynatrace/Splunk (commercial APM), SLI / SLO management (Pyrra / OpenSLO / Sloth — modern SLO-as-code tools), incident response tooling (PagerDuty / Opsgenie / Squadcast / FireHydrant / Rootly / incident.io), chaos engineering (Chaos Monkey / Gremlin / LitmusChaos / Chaos Mesh — for resilience testing), load testing (k6 / Gatling / Locust / JMeter), Terraform/OpenTofu (IaC), ArgoCD/FluxCD (GitOps), service mesh (Istio / Linkerd / Cilium for resilience patterns), distributed systems theory (CAP / consistency models / consensus algorithms basics). Top stack: sre, devops, aws, devsecops, cloud. 71% remote.
The SRE market currently has 188 open roles, 14 of them freshly observed. Median salary $8,615/mo. Observed candidate pool — not published.
71% of SRE jobs are remote or hybrid. SRE work is cloud-based standard. Caveat: on-call rotation requires reliable home-office setup. Time zone overlap critical for distributed SRE teams. International tech companies — full-remote standard. Russian banks — hybrid/office compliance.
⚠ salary known for 6 of 188 jobs; remote share from 173 with a stated format; 14 counted as fresh observations; trend and hiring difficulty are not shown
Demand and observed supply
| Open demand | 188 |
| Observed supply | — not published |
⚠ candidates matching a vacancy are not counted yet: demand and the observed pool are shown
Salary distribution
One in ten earns under $8,510, one in ten over $9,141. Half the market falls between $8,510 and $8,615. Sample: 6.
⚠ percentiles from 6 salaries out of 188 jobs — a guide on a small sample
Demand geography
| country | jobs |
|---|---|
| MX | 20 |
| US | 13 |
| BR | 13 |
| PL | 12 |
| PE | 10 |
| AR | 8 |
| GB | 8 |
| CO | 7 |
| IN | 6 |
| CR | 5 |
The leader by SRE job count is Russia (0 positions). Poland — SRE-friendly EU hub. Germany — Berlin / Munich tech cluster.
⚠ job counts only: salary by country is not published
Used together with
Top SRE stack 2026: Linux + bash mastery (production debugging — perf / bpftrace / flamegraphs), Go (standard for SRE tooling — Kubernetes / Prometheus / etcd) or Python deep, Kubernetes mastery (production-scale), Prometheus + Grafana + Alertmanager + VictoriaMetrics + Mimir (metrics deep), Loki + Tempo + OpenTelemetry (logs + traces), Datadog / New Relic / Dynatrace / Splunk (commercial APM), SLO-as-code (Pyrra / OpenSLO / Sloth), incident response tooling (PagerDuty / Opsgenie / Squadcast / FireHydrant / Rootly / incident.io), chaos engineering (Chaos Monkey / Gremlin / LitmusChaos / Chaos Mesh), load testing (k6 / Gatling / Locust / JMeter), Terraform / OpenTofu, ArgoCD / FluxCD, service mesh (Istio / Linkerd / Cilium), distributed systems theory (CAP / consistency models / consensus algorithms).
Demand by grade
| grade | jobs |
|---|---|
| senior | 22 |
| lead | 9 |
| principal | 4 |
| middle | 1 |
| junior | 1 |
Junior — rare (typical entry DevOps Middle / Backend Middle → SRE Junior). Career flow: DevOps Middle (2-3 years) → SRE Junior (1-2 years) → Middle (2-3 years) → Senior → either Staff / Principal SRE (deep), Engineering Manager (SRE), Backend Distributed Systems Senior pivot, or specialisation in Chaos Engineering / Performance Engineering.
⚠ demand side only: the grade of the observed pool is unknown for most of it
Employers
Demand is spread across 46 employers. The largest accounts for 65.0%, the top ten for 77.1%; the remaining 22.9% is long tail.
| share | value |
|---|---|
| top-1 | 65.0% |
| top-3 | 69.4% |
| top-10 | 77.1% |
| long tail | 22.9% |
⚠ names are not shown: staffing agencies and end employers are not yet told apart by the classifier
Where Zorky sees this market
Observed across 30 sources; the largest accounts for 58.0% — this market does not rest on a single channel.
Recent openings
- Senior Platform SRE - US Shift · IN
- Software Engineer - SRE (Rust) · IN
- Staff Reliability Engineer · US
- DevOps / SRE / DevSecOps Engineer (AWS) - Latin America, Remote
- Site Reliability and Observability Engineer · HK
- Site Reliability Internship - Spring 2027 · US
- SR SRE (Linux & Windows) · BR
- Platform & Reliability Engineer (all genders) · DE
Latest open SRE jobs — the most recent 10 positions with adequate description quality. The full list is in our CRM or via the "see all" link below.
Adjacent markets
SRE overlaps with DevOps (foundation stack), Backend (distributed systems + programming depth), Platform Engineer (internal tooling overlap), Security Engineer (incident response overlap), Performance Engineer (load testing + profile-driven optimisation). Comparison — in the SiblingSubnichesChart above.
⚠ adjacent markets for comparison are not defined yet
How this is measured
- Vacancy
- an open job that cleared the quality gate and lists at least two technologies
- Observed candidate
- a candidate whose stack contains this technology; an aggregate — not a single record leaves the perimeter
- Matchable candidate
- not counted yet
- Window
- jobs open at the moment the snapshot was built
About the data
- Some breakdowns are hidden: their data coverage is not yet sufficient.
- Statistics are shown only where the sample clears a quality gate.
- A missing block does not mean a value of zero.
Breakdowns currently hidden: 8.
Data as of 2026-09-27
Direction: DevOps / SRE
Related specializations
Frequently asked questions
Answers recompute automatically.
What does an SRE Junior, Middle, Senior, or Lead earn?
Junior SRE — rare (typical entry: DevOps Middle / Backend Middle + interest in reliability). The Junior → Middle jump — after the first production incident resolution + the first SLO setup for a service. Middle → Senior — multi-service SLO ownership + on-call mastery (Mean Time to Detect / Recovery metrics ownership) + automation programming at Backend Middle/Senior level. Senior → Staff / Principal — org-wide reliability strategy + production engineering culture leadership + technical mentorship. Career flow: DevOps Middle (2-3 years) → SRE Junior / Middle (1-2 years) → Senior → either Staff / Principal SRE, Engineering Manager (SRE), a move to Backend Distributed Systems Senior, or specialisation in Chaos Engineering or Performance Engineering.
SRE vs DevOps — what's the practical difference (Google distinction + 2026 reality)?
Google's original distinction (Site Reliability Engineering book): DevOps — a culture / philosophy ("break down silos between Dev and Ops"). SRE — a concrete implementation of the DevOps philosophy via specific practices: SLI / SLO / error budgets + toil reduction + automation-first + 50% time on engineering vs operational work + blameless post-mortems. SRE = "what happens when you ask a software engineer to design an operations team". Practical reality 2026: 70% overlap at the stack level (both use K8s + Terraform + Prometheus + Grafana). Differences observable at product companies: 1) Programming depth — SRE writes more custom Go / Python tooling (autoscaling logic, deployment automation, capacity planning algorithms). DevOps Engineer more often configures existing tools. 2) SLI / SLO discipline — SRE owns SLO architecture (Pyrra / OpenSLO / Sloth), error budget policy enforcement, alerting tuned against SLO burn rates. DevOps often sets up monitoring without formal SLO. 3) On-call mastery — SRE on regular on-call rotation (24×7 for critical services), stronger in incident command + post-mortem facilitation. DevOps Engineer is usually on-call for own product only. 4) 50% engineering rule — Google policy: an SRE should not spend >50% time on operational work, the rest is engineering automation. DevOps Engineer has no such guard. 5) Distributed systems theory — SRE interviews often include CAP theorem / consensus algorithms / consistency models / failure mode analysis. DevOps interviews — more practical tooling. In startups this differentiation is often blurred (one person = both DevOps and SRE).
What is the SLI / SLO / error budget framework?
SLI (Service Level Indicator) — a measurable metric of service health (availability / latency / error rate / throughput). SLO (Service Level Objective) — a target value for an SLI (e.g. 99.9% availability over 30 days). SLA (Service Level Agreement) — a contractual commitment to customers (usually weaker than the SLO, e.g. 99.5% if SLO is 99.9%, to keep a safety margin). Error budget = 100% − SLO. If the SLO is 99.9% over 30 days → error budget = 0.1% = 43 minutes of downtime/month. When the error budget burns fast → freeze new feature deployments, focus the team on reliability work. When the error budget is healthy → ship features aggressively. Practical framework setup (12 steps): 1) Identify customer-facing critical user journeys (CUJs). 2) Pick SLIs for each CUJ (typically availability + latency for synchronous, throughput + freshness for async). 3) Choose the initial SLO target (rule: slightly below current performance). 4) Set up SLI measurement (Prometheus + Grafana or managed). 5) Configure burn-rate alerts (multi-window: fast burn 1h 14.4× rate, slow burn 6h 6× rate — Google formula). 6) Set up SLO-as-code (Pyrra / OpenSLO / Sloth) for version control. 7) Document the error budget policy (what happens at exhaustion — feature freeze? incident review?). 8) Quarterly SLO review (target adjustment based on actual performance + customer impact). 9) Toil tracking + reduction roadmap (target: <50% time on toil). 10) Post-mortem culture — blameless, focus on action items. 11) Chaos engineering integration (Gremlin / Chaos Mesh — pre-test SLO under failure). 12) Customer trust dashboard (public status page — Statuspage / Atlassian / Better Uptime).
What skills and tools does an SRE need?
Linux / systems deep: processes, namespaces, cgroups, networking (TCP/IP, DNS, load balancing), performance debugging (strace / perf / eBPF). Observability stack: Prometheus + Grafana (PromQL deep), distributed tracing (Tempo / Jaeger / OpenTelemetry), logs (Loki / ELK), managed APM (Datadog / Grafana Cloud / Honeycomb) — SRE "sees" the system through metrics. SLO practice: SLI / SLO / error budget (see separate question), burn-rate alerts, SLO-as-code. Incident response + on-call: PagerDuty / Opsgenie, runbooks, blameless post-mortems, severity assessment. Kubernetes: production operations (CKA level) — workloads, networking, troubleshooting. IaC + automation: Terraform / OpenTofu, Ansible; the key SRE skill is reducing toil (manual repeated work) via automation, target <50% of time on toil. CI/CD: safe deploys — canary, blue-green, progressive delivery, rollback. Programming: Python and / or Go at a level of writing maintainable automation tools, not "scripts". Distributed systems: failure modes, retry / timeout / circuit breaker, idempotency, consistency — the foundation for capacity planning and DR. Chaos engineering: Gremlin / Chaos Mesh — verify reliability before the incident. The main point: SRE treats reliability as a product — measures it (SLO), automates routine, and systematically removes causes of incidents instead of putting them out by hand. English is mandatory — the SRE literature (Google SRE book) and the community are English-speaking.
Can SREs work remotely?
Yes, 71% of SRE jobs are full-remote or hybrid. SRE work — cloud-based + monitoring dashboards. Outsourcers — almost always remote. Russian product companies — hybrid or remote after probation. Russian banks — hybrid/office security compliance. International tech companies — full-remote standard. Caveat for SRE specifically: on-call rotation — requires reliable internet + power backup + quiet space for night-emergency response. Some companies require a home-office setup audit before a remote SRE offer. Time zone — SRE roles usually require overlap with the team's primary timezone (US companies often want 4+ hours overlap with PT/ET). Relocant hubs: Poland / Germany / Canada / Serbia / Georgia. English for international SRE remote — must (incident command on Zoom in English under stress).
How is Production Engineer (Facebook / Meta term) different from SRE?
Production Engineer (PE) — Facebook / Meta's term for SRE. Same discipline, almost identical responsibilities — focus reliability + automation + on-call + capacity planning + distributed systems. The difference is historically philosophical: Google SRE — "a software engineer who happens to do ops", Facebook PE — "an engineer embedded in a product team for reliability". In 2026 practice — almost fully overlapping. Other equivalent titles: Reliability Engineer (LinkedIn), Infrastructure Engineer (often overlaps with SRE), Production Operations Engineer (legacy term). How to read job postings 2026: look for signals in the JD — if it mentions "SLI/SLO", "error budgets", "toil reduction", "50% engineering time", "blameless post-mortems", "on-call rotation" — it's an SRE-style role regardless of title. If it mentions "CI/CD setup", "cloud migration", "infrastructure as code" but WITHOUT SLO mentions — it's general DevOps.
Where to start in SRE in 2026?
Roadmap: 1) DevOps base solid — Linux mastery + Docker + Kubernetes (CKA) + cloud platform deeply + IaC (Terraform). Without this base there's no point going into SRE. 2) Programming Backend Middle level — Go (standard for SRE — Kubernetes / Prometheus / most SRE tooling in Go) or Python deep (data analysis + scripting). Books: "The Go Programming Language" Donovan / Kernighan, "Fluent Python" Ramalho. 3) "Site Reliability Engineering" Google book (free PDF) — must-read, read twice. 4) "The Site Reliability Workbook" Google — practical complement (case studies + exercises). 5) SLI / SLO mastery — set up SLO-as-code (Sloth / Pyrra / OpenSLO) for a real service, configure burn-rate alerts (multi-window: fast 1h 14.4×, slow 6h 6×). 6) Distributed systems theory — CAP theorem, consistency models (linearizability / sequential / causal / eventual), consensus (Paxos / Raft basics), failure mode analysis. Books: "Designing Data-Intensive Applications" Martin Kleppmann (must-read 2026), "Database Internals" Petrov. 7) Chaos engineering — Chaos Mesh / Gremlin / LitmusChaos. Set up chaos experiments on own K8s cluster. Book: "Chaos Engineering" Nora Jones / Casey Rosenthal. 8) Load testing mastery — k6 (modern JS-based, rising) or Gatling (Scala DSL) or Locust (Python). Set up load tests integrated into CI. 9) Observability deep: Prometheus advanced (PromQL mastery + recording rules + federation), Grafana advanced (templating + alerting), Loki + Tempo + OpenTelemetry. Use case: distributed tracing across microservices. 10) Incident response training — incident command basics, blameless post-mortems framework, communication during incidents. Resources: "Incident Response & Computer Forensics" Luttgens / Pepe / Mandia, Google's incident management training. 11) Pet project: deploy a distributed app on K8s with full SLO setup + chaos experiments + on-call simulation. Document as a production-ready system. RU courses: Slurm SRE, Otus "SRE", Karpov.Courses SRE Track. International (EN): "Database Reliability Engineering" Campbell / Majors, USENIX SREcon talks (free YouTube), Google Cloud SRE Certification Path. DevOps Middle + interest → SRE Junior — 4-10 months (need to strengthen programming + distributed systems theory).
How many SRE jobs are open across CIS and Europe?
188 active open SRE positions — premium segment of the DevOps direction. Geography: Russia / Poland / remote. The real market is broader thanks to a huge international remote segment. Time to close a Senior SRE role — 6-12 weeks (longer than general DevOps due to rare-skill requirements: programming + distributed systems + on-call mastery combination).
What skills does a Senior SRE need?
A Senior SRE owns the full reliability engineering cycle + technical leadership. Programming Backend Middle+ level: Go mastery (standard for SRE automation) or Python deep — at a level of "can write production-grade autoscaling logic / capacity planning algorithms / custom K8s operators". Kubernetes mastery deep: production-scale (nodes), Operators (Kubebuilder / Operator SDK), custom CRDs, multi-tenancy patterns. Distributed systems theory: deep CAP theorem understanding, consistency models (linearizability / sequential / causal / eventual), consensus algorithms (Paxos / Raft), failure mode taxonomy, network partition handling patterns. Observability mastery: Prometheus advanced (PromQL mastery + recording rules + federation + remote_write), Grafana advanced (templating + transformations + alerting), Loki + Tempo + OpenTelemetry integration mastery, distributed tracing across microservices. SLI / SLO architecture mastery: SLO-as-code (Pyrra / OpenSLO / Sloth), error budget policy design + enforcement automation, multi-window burn-rate alert tuning (avoid alert fatigue). Incident command mastery: lead incident response under stress, blameless post-mortem facilitation, contributing factors analysis (not root cause — modern thinking), action items prioritisation. Chaos engineering mastery: Chaos Mesh / Gremlin / LitmusChaos — design chaos experiments, GameDay facilitation. Capacity planning mastery: load testing methodology (k6 / Gatling / Locust), resource forecasting models, headroom analysis, peak-load handling. Performance engineering: profile-driven optimisation (perf / bpftrace / flamegraphs), memory leak diagnosis, GC tuning, network performance analysis. Service mesh deep: Istio / Linkerd / Cilium — traffic management, retry policies, circuit breakers, fault injection. System design for reliability: design multi-region multi-AZ HA on whiteboard, RPO / RTO planning, DR strategies, cell-based architecture for blast radius limitation. Soft: ADRs writing, incident communications (status page updates + stakeholder calls in crisis), on-call rotation discipline, cross-team collaboration (Backend / DevOps / Platform / Security teams), mentoring Middle SRE. English for Senior+ MUST — SRE is intensely cross-team + the community is English-speaking (USENIX SREcon, papers).
Leave a request
Describe the task and leave a contact — the request goes to our CRM and we reply at the contact you provide.