Zorky CRMZorky CRM
EN|RU
@termdocs

SRE: market

Site Reliability Engineer (SRE) — premium role inside the DevOps direction, invented by Google in 2003. Focus: reliability + SLI / SLO / error budgets + incident response + automation to reduce toil. Programming-heavier than general DevOps (Go / Python for automation + custom tooling). Role family: SRE (mid — owns reliability of one service), Senior SRE (multi-service + on-call mastery + SLO architecture), Staff / Principal SRE (org-wide reliability strategy + production engineering culture leadership), SRE Tech Lead (team + reliability roadmap), Production Engineer (alternative title — Facebook / Meta term). Stack 2026: Linux+bash deep mastery (production debugging), Go (standard for SRE automation — Kubernetes + Prometheus + most SRE tooling in Go), Python (data analysis + scripting), Kubernetes mastery (production-scale), Prometheus+Grafana+Alertmanager+VictoriaMetrics+Mimir mastery (metrics deep), Loki+Tempo+OpenTelemetry (logs + traces), Datadog/New Relic/Dynatrace/Splunk (commercial APM), SLI / SLO management (Pyrra / OpenSLO / Sloth — modern SLO-as-code tools), incident response tooling (PagerDuty / Opsgenie / Squadcast / FireHydrant / Rootly / incident.io), chaos engineering (Chaos Monkey / Gremlin / LitmusChaos / Chaos Mesh — for resilience testing), load testing (k6 / Gatling / Locust / JMeter), Terraform/OpenTofu (IaC), ArgoCD/FluxCD (GitOps), service mesh (Istio / Linkerd / Cilium for resilience patterns), distributed systems theory (CAP / consistency models / consensus algorithms basics). Top stack: sre, devops, aws, devsecops, cloud. 71% remote.

188
open jobs
$8,615
median $/mo
—
observed supply
77%
remote

The SRE market currently has 188 open roles, 14 of them freshly observed. Median salary $8,615/mo. Observed candidate pool — not published.

71% of SRE jobs are remote or hybrid. SRE work is cloud-based standard. Caveat: on-call rotation requires reliable home-office setup. Time zone overlap critical for distributed SRE teams. International tech companies — full-remote standard. Russian banks — hybrid/office compliance.

⚠ salary known for 6 of 188 jobs; remote share from 173 with a stated format; 14 counted as fresh observations; trend and hiring difficulty are not shown

Demand and observed supply

Open demand188
Observed supply— not published

⚠ candidates matching a vacancy are not counted yet: demand and the observed pool are shown

Salary distribution

One in ten earns under $8,510, one in ten over $9,141. Half the market falls between $8,510 and $8,615. Sample: 6.

⚠ percentiles from 6 salaries out of 188 jobs — a guide on a small sample

Demand geography

countryjobs
MX20
US13
BR13
PL12
PE10
AR8
GB8
CO7
IN6
CR5

The leader by SRE job count is Russia (0 positions). Poland — SRE-friendly EU hub. Germany — Berlin / Munich tech cluster.

⚠ job counts only: salary by country is not published

Used together with

sre 148aws 129devsecops 122cloud 38python 15site reliability 15kubernetes 13docker 9azure 8go 8gcp 6jenkins 6

Top SRE stack 2026: Linux + bash mastery (production debugging — perf / bpftrace / flamegraphs), Go (standard for SRE tooling — Kubernetes / Prometheus / etcd) or Python deep, Kubernetes mastery (production-scale), Prometheus + Grafana + Alertmanager + VictoriaMetrics + Mimir (metrics deep), Loki + Tempo + OpenTelemetry (logs + traces), Datadog / New Relic / Dynatrace / Splunk (commercial APM), SLO-as-code (Pyrra / OpenSLO / Sloth), incident response tooling (PagerDuty / Opsgenie / Squadcast / FireHydrant / Rootly / incident.io), chaos engineering (Chaos Monkey / Gremlin / LitmusChaos / Chaos Mesh), load testing (k6 / Gatling / Locust / JMeter), Terraform / OpenTofu, ArgoCD / FluxCD, service mesh (Istio / Linkerd / Cilium), distributed systems theory (CAP / consistency models / consensus algorithms).

Demand by grade

gradejobs
senior22
lead9
principal4
middle1
junior1

Junior — rare (typical entry DevOps Middle / Backend Middle → SRE Junior). Career flow: DevOps Middle (2-3 years) → SRE Junior (1-2 years) → Middle (2-3 years) → Senior → either Staff / Principal SRE (deep), Engineering Manager (SRE), Backend Distributed Systems Senior pivot, or specialisation in Chaos Engineering / Performance Engineering.

⚠ demand side only: the grade of the observed pool is unknown for most of it

Employers

Demand is spread across 46 employers. The largest accounts for 65.0%, the top ten for 77.1%; the remaining 22.9% is long tail.

sharevalue
top-165.0%
top-369.4%
top-1077.1%
long tail22.9%

⚠ names are not shown: staffing agencies and end employers are not yet told apart by the classifier

Where Zorky sees this market

Observed across 30 sources; the largest accounts for 58.0% — this market does not rest on a single channel.

Recent openings

All jobs →

Latest open SRE jobs — the most recent 10 positions with adequate description quality. The full list is in our CRM or via the "see all" link below.

Adjacent markets

DevOps / SREDevOps EngineerPlatform EngineerCloud EngineerKubernetes / ContainerInfrastructure EngineerDevSecOps

SRE overlaps with DevOps (foundation stack), Backend (distributed systems + programming depth), Platform Engineer (internal tooling overlap), Security Engineer (incident response overlap), Performance Engineer (load testing + profile-driven optimisation). Comparison — in the SiblingSubnichesChart above.

⚠ adjacent markets for comparison are not defined yet

How this is measured
Vacancy
an open job that cleared the quality gate and lists at least two technologies
Observed candidate
a candidate whose stack contains this technology; an aggregate — not a single record leaves the perimeter
Matchable candidate
not counted yet
Window
jobs open at the moment the snapshot was built

About the data

  • Some breakdowns are hidden: their data coverage is not yet sufficient.
  • Statistics are shown only where the sample clears a quality gate.
  • A missing block does not mean a value of zero.

Breakdowns currently hidden: 8.

Data as of 2026-09-27

Direction: DevOps / SRE

Related specializations

Cloud EngineerDevSecOpsInfrastructure EngineerKubernetes / ContainerPlatform Engineer

Frequently asked questions

Answers recompute automatically.

What does an SRE Junior, Middle, Senior, or Lead earn?

Junior SRE — rare (typical entry: DevOps Middle / Backend Middle + interest in reliability). The Junior → Middle jump — after the first production incident resolution + the first SLO setup for a service. Middle → Senior — multi-service SLO ownership + on-call mastery (Mean Time to Detect / Recovery metrics ownership) + automation programming at Backend Middle/Senior level. Senior → Staff / Principal — org-wide reliability strategy + production engineering culture leadership + technical mentorship. Career flow: DevOps Middle (2-3 years) → SRE Junior / Middle (1-2 years) → Senior → either Staff / Principal SRE, Engineering Manager (SRE), a move to Backend Distributed Systems Senior, or specialisation in Chaos Engineering or Performance Engineering.

SRE vs DevOps — what's the practical difference (Google distinction + 2026 reality)?

Google's original distinction (Site Reliability Engineering book): DevOps — a culture / philosophy ("break down silos between Dev and Ops"). SRE — a concrete implementation of the DevOps philosophy via specific practices: SLI / SLO / error budgets + toil reduction + automation-first + 50% time on engineering vs operational work + blameless post-mortems. SRE = "what happens when you ask a software engineer to design an operations team". Practical reality 2026: 70% overlap at the stack level (both use K8s + Terraform + Prometheus + Grafana). Differences observable at product companies: 1) Programming depth — SRE writes more custom Go / Python tooling (autoscaling logic, deployment automation, capacity planning algorithms). DevOps Engineer more often configures existing tools. 2) SLI / SLO discipline — SRE owns SLO architecture (Pyrra / OpenSLO / Sloth), error budget policy enforcement, alerting tuned against SLO burn rates. DevOps often sets up monitoring without formal SLO. 3) On-call mastery — SRE on regular on-call rotation (24×7 for critical services), stronger in incident command + post-mortem facilitation. DevOps Engineer is usually on-call for own product only. 4) 50% engineering rule — Google policy: an SRE should not spend >50% time on operational work, the rest is engineering automation. DevOps Engineer has no such guard. 5) Distributed systems theory — SRE interviews often include CAP theorem / consensus algorithms / consistency models / failure mode analysis. DevOps interviews — more practical tooling. In startups this differentiation is often blurred (one person = both DevOps and SRE).

What is the SLI / SLO / error budget framework?

SLI (Service Level Indicator) — a measurable metric of service health (availability / latency / error rate / throughput). SLO (Service Level Objective) — a target value for an SLI (e.g. 99.9% availability over 30 days). SLA (Service Level Agreement) — a contractual commitment to customers (usually weaker than the SLO, e.g. 99.5% if SLO is 99.9%, to keep a safety margin). Error budget = 100% − SLO. If the SLO is 99.9% over 30 days → error budget = 0.1% = 43 minutes of downtime/month. When the error budget burns fast → freeze new feature deployments, focus the team on reliability work. When the error budget is healthy → ship features aggressively. Practical framework setup (12 steps): 1) Identify customer-facing critical user journeys (CUJs). 2) Pick SLIs for each CUJ (typically availability + latency for synchronous, throughput + freshness for async). 3) Choose the initial SLO target (rule: slightly below current performance). 4) Set up SLI measurement (Prometheus + Grafana or managed). 5) Configure burn-rate alerts (multi-window: fast burn 1h 14.4× rate, slow burn 6h 6× rate — Google formula). 6) Set up SLO-as-code (Pyrra / OpenSLO / Sloth) for version control. 7) Document the error budget policy (what happens at exhaustion — feature freeze? incident review?). 8) Quarterly SLO review (target adjustment based on actual performance + customer impact). 9) Toil tracking + reduction roadmap (target: <50% time on toil). 10) Post-mortem culture — blameless, focus on action items. 11) Chaos engineering integration (Gremlin / Chaos Mesh — pre-test SLO under failure). 12) Customer trust dashboard (public status page — Statuspage / Atlassian / Better Uptime).

What skills and tools does an SRE need?

Linux / systems deep: processes, namespaces, cgroups, networking (TCP/IP, DNS, load balancing), performance debugging (strace / perf / eBPF). Observability stack: Prometheus + Grafana (PromQL deep), distributed tracing (Tempo / Jaeger / OpenTelemetry), logs (Loki / ELK), managed APM (Datadog / Grafana Cloud / Honeycomb) — SRE "sees" the system through metrics. SLO practice: SLI / SLO / error budget (see separate question), burn-rate alerts, SLO-as-code. Incident response + on-call: PagerDuty / Opsgenie, runbooks, blameless post-mortems, severity assessment. Kubernetes: production operations (CKA level) — workloads, networking, troubleshooting. IaC + automation: Terraform / OpenTofu, Ansible; the key SRE skill is reducing toil (manual repeated work) via automation, target <50% of time on toil. CI/CD: safe deploys — canary, blue-green, progressive delivery, rollback. Programming: Python and / or Go at a level of writing maintainable automation tools, not "scripts". Distributed systems: failure modes, retry / timeout / circuit breaker, idempotency, consistency — the foundation for capacity planning and DR. Chaos engineering: Gremlin / Chaos Mesh — verify reliability before the incident. The main point: SRE treats reliability as a product — measures it (SLO), automates routine, and systematically removes causes of incidents instead of putting them out by hand. English is mandatory — the SRE literature (Google SRE book) and the community are English-speaking.

Can SREs work remotely?

Yes, 71% of SRE jobs are full-remote or hybrid. SRE work — cloud-based + monitoring dashboards. Outsourcers — almost always remote. Russian product companies — hybrid or remote after probation. Russian banks — hybrid/office security compliance. International tech companies — full-remote standard. Caveat for SRE specifically: on-call rotation — requires reliable internet + power backup + quiet space for night-emergency response. Some companies require a home-office setup audit before a remote SRE offer. Time zone — SRE roles usually require overlap with the team's primary timezone (US companies often want 4+ hours overlap with PT/ET). Relocant hubs: Poland / Germany / Canada / Serbia / Georgia. English for international SRE remote — must (incident command on Zoom in English under stress).

How is Production Engineer (Facebook / Meta term) different from SRE?

Production Engineer (PE) — Facebook / Meta's term for SRE. Same discipline, almost identical responsibilities — focus reliability + automation + on-call + capacity planning + distributed systems. The difference is historically philosophical: Google SRE — "a software engineer who happens to do ops", Facebook PE — "an engineer embedded in a product team for reliability". In 2026 practice — almost fully overlapping. Other equivalent titles: Reliability Engineer (LinkedIn), Infrastructure Engineer (often overlaps with SRE), Production Operations Engineer (legacy term). How to read job postings 2026: look for signals in the JD — if it mentions "SLI/SLO", "error budgets", "toil reduction", "50% engineering time", "blameless post-mortems", "on-call rotation" — it's an SRE-style role regardless of title. If it mentions "CI/CD setup", "cloud migration", "infrastructure as code" but WITHOUT SLO mentions — it's general DevOps.

Where to start in SRE in 2026?

Roadmap: 1) DevOps base solid — Linux mastery + Docker + Kubernetes (CKA) + cloud platform deeply + IaC (Terraform). Without this base there's no point going into SRE. 2) Programming Backend Middle level — Go (standard for SRE — Kubernetes / Prometheus / most SRE tooling in Go) or Python deep (data analysis + scripting). Books: "The Go Programming Language" Donovan / Kernighan, "Fluent Python" Ramalho. 3) "Site Reliability Engineering" Google book (free PDF) — must-read, read twice. 4) "The Site Reliability Workbook" Google — practical complement (case studies + exercises). 5) SLI / SLO mastery — set up SLO-as-code (Sloth / Pyrra / OpenSLO) for a real service, configure burn-rate alerts (multi-window: fast 1h 14.4×, slow 6h 6×). 6) Distributed systems theory — CAP theorem, consistency models (linearizability / sequential / causal / eventual), consensus (Paxos / Raft basics), failure mode analysis. Books: "Designing Data-Intensive Applications" Martin Kleppmann (must-read 2026), "Database Internals" Petrov. 7) Chaos engineering — Chaos Mesh / Gremlin / LitmusChaos. Set up chaos experiments on own K8s cluster. Book: "Chaos Engineering" Nora Jones / Casey Rosenthal. 8) Load testing mastery — k6 (modern JS-based, rising) or Gatling (Scala DSL) or Locust (Python). Set up load tests integrated into CI. 9) Observability deep: Prometheus advanced (PromQL mastery + recording rules + federation), Grafana advanced (templating + alerting), Loki + Tempo + OpenTelemetry. Use case: distributed tracing across microservices. 10) Incident response training — incident command basics, blameless post-mortems framework, communication during incidents. Resources: "Incident Response & Computer Forensics" Luttgens / Pepe / Mandia, Google's incident management training. 11) Pet project: deploy a distributed app on K8s with full SLO setup + chaos experiments + on-call simulation. Document as a production-ready system. RU courses: Slurm SRE, Otus "SRE", Karpov.Courses SRE Track. International (EN): "Database Reliability Engineering" Campbell / Majors, USENIX SREcon talks (free YouTube), Google Cloud SRE Certification Path. DevOps Middle + interest → SRE Junior — 4-10 months (need to strengthen programming + distributed systems theory).

How many SRE jobs are open across CIS and Europe?

188 active open SRE positions — premium segment of the DevOps direction. Geography: Russia / Poland / remote. The real market is broader thanks to a huge international remote segment. Time to close a Senior SRE role — 6-12 weeks (longer than general DevOps due to rare-skill requirements: programming + distributed systems + on-call mastery combination).

What skills does a Senior SRE need?

A Senior SRE owns the full reliability engineering cycle + technical leadership. Programming Backend Middle+ level: Go mastery (standard for SRE automation) or Python deep — at a level of "can write production-grade autoscaling logic / capacity planning algorithms / custom K8s operators". Kubernetes mastery deep: production-scale (nodes), Operators (Kubebuilder / Operator SDK), custom CRDs, multi-tenancy patterns. Distributed systems theory: deep CAP theorem understanding, consistency models (linearizability / sequential / causal / eventual), consensus algorithms (Paxos / Raft), failure mode taxonomy, network partition handling patterns. Observability mastery: Prometheus advanced (PromQL mastery + recording rules + federation + remote_write), Grafana advanced (templating + transformations + alerting), Loki + Tempo + OpenTelemetry integration mastery, distributed tracing across microservices. SLI / SLO architecture mastery: SLO-as-code (Pyrra / OpenSLO / Sloth), error budget policy design + enforcement automation, multi-window burn-rate alert tuning (avoid alert fatigue). Incident command mastery: lead incident response under stress, blameless post-mortem facilitation, contributing factors analysis (not root cause — modern thinking), action items prioritisation. Chaos engineering mastery: Chaos Mesh / Gremlin / LitmusChaos — design chaos experiments, GameDay facilitation. Capacity planning mastery: load testing methodology (k6 / Gatling / Locust), resource forecasting models, headroom analysis, peak-load handling. Performance engineering: profile-driven optimisation (perf / bpftrace / flamegraphs), memory leak diagnosis, GC tuning, network performance analysis. Service mesh deep: Istio / Linkerd / Cilium — traffic management, retry policies, circuit breakers, fault injection. System design for reliability: design multi-region multi-AZ HA on whiteboard, RPO / RTO planning, DR strategies, cell-based architecture for blast radius limitation. Soft: ADRs writing, incident communications (status page updates + stakeholder calls in crisis), on-call rotation discipline, cross-team collaboration (Backend / DevOps / Platform / Security teams), mentoring Middle SRE. English for Senior+ MUST — SRE is intensely cross-team + the community is English-speaking (USENIX SREcon, papers).

Leave a request

Describe the task and leave a contact — the request goes to our CRM and we reply at the contact you provide.