Description
SRE Lead / Senior Site Reliability Engineer
Company: Zenith
Zenith is the first Ethereum extension for the Canton Network, enabling fully Ethereum-compatible execution environments that integrate with Canton's privacy, compliance, and settlement infrastructure. Zenith EVM and Zenith Stack let teams and institutions deploy and run Ethereum applications and customized blockchain environments engineered for high-frequency institutional finance.
Applicants based in Europe are strongly preferred.
Responsibilities:
- Own production reliability across Zenith’s internal systems, Zenith Stack environments, and validator infrastructure.
- Manage and evolve cloud environments, building automation and scalable systems for multi-layered blockchain execution and off-chain runtimes.
- Define and implement incident response, observability, monitoring, alerting, on-call rotations, runbooks, and reliability KPIs.
- Establish versioning, release management, configuration, secrets handling, environment orchestration, and reproducibility practices.
- Support deployment, scaling, and reliability of client-tailored Zenith Stack environments.
- Partner with core engineering on validator operations, network participation, upgrades, and secure infrastructure.
- Drive automation to eliminate manual inefficiencies in provisioning, deployment, and system testing.
Requirements:
- 7+ years in site reliability engineering, DevOps, platform engineering, or infrastructure-focused software engineering.
- Proven experience owning reliability, uptime, and operational excellence for distributed or mission-critical systems.
- Prior experience building or leading an SRE team; able to run systems solo initially and scale processes as the team grows.
- Deep understanding of Linux systems, networking, cloud platforms (AWS/GCP/Azure), and containerization.
- Kubernetes expertise (non-negotiable).
- Hands-on experience building monitoring, observability, performance, alerting, incident response systems, and on-call processes.
- Familiarity with Infrastructure-as-Code (Terraform, Pulumi), CI/CD, secrets management, and logging & metrics stacks.
Nice to have:
- Experience operating infrastructure for blockchain networks, validators, L1/L2 systems, or consensus clients.
- Familiarity with EVM-based blockchain architectures.
- Experience with high-throughput, low-latency, or cryptographic workloads.
- Background supporting enterprise clients, regulated environments, or high-availability financial systems.
Stack:
Linux, Kubernetes, Docker, AWS/GCP/Azure, Terraform/Pulumi, CI/CD, secrets management, logging & metrics
Benefits:
- Remote-first with in-person meetups 2–4 times/year.
- Culture focused on work-life fit, autonomy, and sustainable performance.
📩 Apply
Employer contacts (email/phone/telegram) are hidden from the public preview —
send your CV, and we will connect you directly.