Описание
Site Reliability Engineer — Telemetry
Company: Payward (Kraken)
Payward, the parent company of Kraken, builds modern, globally accessible financial infrastructure to advance an open financial system. Kraken is a long-standing crypto platform trusted by over 10 million individuals and institutions, offering spot trading, margin, futures, staking, and OTC services.
Responsibilities:
- Operate and improve the shared telemetry platform for metrics, logs, traces, alerting, dashboards, and profiling.
- Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems and VictoriaMetrics/Grafana.
- Operate log pipelines (Vector, Splunk, Loki) and distributed tracing/profiling (Grafana Alloy, Tempo, OpenTelemetry, Pyroscope).
- Deploy and manage telemetry services with Terraform, Terragrunt, and container orchestration across environments.
- Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
- Build reusable configuration and automation for dashboards, alerts, and telemetry integrations.
- Participate in incident response and on-call, write runbooks, and iterate on platform reliability.
Requirements:
- 3+ years as an SRE, Platform/Infrastructure Engineer, Observability Engineer, or similar.
- Experience managing production systems that collect, process, store, and serve telemetry (metrics, logs, traces, profiles).
- Experience with Prometheus or a Prometheus-compatible monitoring stack (collection, querying, alerting).
- Experience troubleshooting distributed production systems (availability, latency, data flow, capacity).
- Experience with Infrastructure as Code (Terraform) and CI/CD.
- Experience operating containerised workloads with Nomad, Kubernetes, or similar.
- Solid scripting/programming skills and comfort using AI tools/agents to accelerate delivery.
Nice to have:
- Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
- PromQL or LogQL, dashboards/alerts as code, and maintaining operators/CRDs in Kubernetes.
- Experience with Consul, Vault, AWS, on-prem/datacentre infrastructure, or high-volume logging/streaming pipelines.
- Background in regulated/financial services environments (change management, audit trails).
Stack:
Prometheus, VictoriaMetrics, Grafana, Grafana Alloy, Tempo, Loki, Vector, Splunk, OpenTelemetry, Pyroscope, Terraform, Terragrunt, Nomad, Kubernetes, PromQL, LogQL, Consul, Vault, AWS
Benefits:
- Compensation: $120k - $131k (estimated)
📩 Apply
Контакты работодателя (email/phone/telegram) скрыты из публичного превью —
отправьте резюме, чтобы мы связали вас напрямую.