Описание
Senior Site Reliability Engineer / SRE Lead
Company: Zenith
Zenith is the first Ethereum extension for the Canton Network, enabling teams to run Ethereum-compatible execution environments integrated with Canton's privacy, compliance, and settlement infrastructure. Their Zenith EVM and Zenith Stack let builders and institutions deploy high-throughput, low-latency blockchain environments using familiar tooling while benefiting from institutional-grade privacy and settlement. The company focuses on bridging DeFi and TradFi, operating validator infrastructure and supporting institutional-grade ledger deployments.
Candidates based in Europe are strongly preferred.
Responsibilities:
- Own production reliability and ensure operational resilience across internal systems, Zenith Stack environments, and validator infrastructure.
- Manage and evolve scalable cloud environments, automation, and infrastructure supporting multi-layered blockchain execution.
- Define and implement incident response, observability, monitoring, alerting, on-call rotation, runbooks, and reliability KPIs.
- Lead versioning, release management, configuration, secrets, environment orchestration, and reproducibility practices.
- Support client Zenith Stack deployments, scaling, and reliability for client-tailored environments.
- Collaborate on validator and protocol operations, upgrades, and secure infrastructure.
- Drive automation to replace manual processes for provisioning, deployment, and system testing.
Requirements:
- 7+ years in site reliability engineering, DevOps, platform engineering, or infrastructure-focused software engineering.
- Proven track record owning reliability, uptime, and operational excellence for distributed or mission-critical systems.
- Experience building or leading an SRE team; able to operate solo initially and grow a team over time.
- Deep knowledge of Linux systems, networking, containerization, and cloud platforms (AWS/GCP/Azure).
- Kubernetes expertise (non-negotiable) and hands-on experience with monitoring, observability, alerting, and incident response.
- Experience with Infrastructure-as-Code (Terraform, Pulumi), CI/CD systems, secrets management, and logging & metrics stacks.
Nice to have:
- Experience operating infrastructure for blockchain networks, validators, L1/L2 systems, or consensus clients.
- Familiarity with EVM-based architectures and deploying high-throughput, low-latency or cryptographic workloads.
- Background supporting enterprise clients, regulated environments, or high-availability financial systems.
Stack:
Linux, AWS/GCP/Azure, Kubernetes, Terraform/Pulumi, CI/CD, Secrets management, Logging & metrics, Containerization
Benefits:
- Remote-first team with 2–4 in-person meetings per year.
- Emphasis on autonomy, ownership, work-life fit, and sustainable performance.
- Culture focused on trust, feedback, and operating with precision.
📩 Apply
Контакты работодателя (email/phone/telegram) скрыты из публичного превью —
отправьте резюме, чтобы мы связали вас напрямую.