Описание
Key
Responsibilities
Incident & SLA Management Respond to, triage, and resolve incidents (P1–P4) within contracted SLAs, owning each ticket from detection through to resolution. Drive root-cause analysis (RCA) and contribute to post-incident reviews for major incidents. Provide clear, timely updates to customers and internal stakeholders throughout the incident lifecycle.
Monitoring, Alerting & Standby Serve on the standby / on-call rota and respond promptly to alerts raised by the monitoring platform (Elastic APM and associated observability tooling). Proactively monitor dashboards and system health to detect and address issues before they impact service. Tune alert thresholds and reduce alert noise to improve signal quality and response efficiency.
Troubleshooting & Cross-TeamCollaboration Partner with application, network, security, platform, and other engineering teams to diagnose and resolve complex issues spanning multiple domains. , AWS Support) where required, and manage the escalation through to closure. Operational Support (BAU) Perform day-to-day operational support of AWS environments — compute (EC2), storage (S3 / EBS), networking (VPC, security groups, load balancers), IAM, patching, and backup/restore.
Execute changes in line with ITIL change management processes and maintain a strong compliance and audit posture. Contribute to problem management and continuous service-improvement initiatives. Documentation & Governance Maintain and improve runbooks, standard operating procedures (SOPs), and knowledge-base articles.
, IM8) and GCC (Government Commercial Cloud) governance and operating standards. Required
Qualifications
& Experience Bachelor’s degree in Computer Science, Information Technology, or related field. 3–4 years of relevant experience as a Day-2 Cloud / Cloud Operations Engineer supporting production AWS environments. Solid working knowledge of core AWS services (EC2, S3, VPC, IAM, CloudWatch, ELB, RDS).
Hands-on experience with ITIL-aligned incident, problem, and change management in a managed-services or operations context. Demonstrated experience supporting Singapore Government accounts. GCC (Government Commercial Cloud) experience — familiarity with GCC operating models, controls, and governance.
Willingness and ability to participate in a standby / on-call and shift rota. Strong troubleshooting mindset with the ability to remain composed and effective during high-pressure incidents. Preferred (Nice-to-Have) Experience with Elastic APM / the Elastic Stack (Elasticsearch, Kibana) for application performance monitoring and observability.
AWS certification — AWS Certified SysOps Administrator – Associate or Solutions Architect – Associate. Scripting / automation skills (Python, Bash) and Infrastructure-as-Code (Terraform or CloudFormation). ITIL Foundation certification.
Exposure to other observability / monitoring tools (CloudWatch, Grafana, Prometheus).
Контакты работодателя (email/phone/telegram) скрыты из публичного превью —
отправьте резюме, чтобы мы связали вас напрямую.