Description
Responsibilities: Build and maintain observability using logs, metrics, and dashboards for on-prem and cloud systems. Drive reliability improvements with performance analysis, capacity planning, stress tests, restore drills, failure prevention. Fine-tune alerts to reduce noise and improve signal quality and response speed; validate monitoring/alerts after releases and during deployments.
Define and track reliability targets (SLA/RTO) for critical services. Improve incident response by maintaining operational procedures, service catalogs and clear escalation paths. g.
health checks, remediation, validation gates. Perform Root Cause Analysis for reliability incidents and implement preventative actions. Ensure observability configuration changes are controlled and audit-evidenced.
Uphold IT General Control and compliance standard with evidence retention, access controls, change approvals.
Requirements
Computer Science or related Engineering Degree (or above) 3 years or more working experience in SRE/DevOps duties
Employer contacts (email/phone/telegram) are hidden from the public preview —
send your CV, and we will connect you directly.