Zorky CRMZorky CRM
EN|RU
@termdocs
← All jobs

Senior Data Engineer

principalhybrid8,000 SGDSingapore, SGScore 75/1001d ago
Stack
DashboardsApplicationsScalabilitySecurityPlatform ModernizationMonitoring SolutionsOperational Delivery TeamsOperations and MaintenanceLoggingTelemetryDeployment AutomationHybrid Cloud
Apply
Upload your CV — we will connect you with the employer directly through our pool.
Send your CV →
Description
Activate Interactive Pte Ltd (“Activate”) is a leading technology consultancy headquartered in Singapore with a presence in Malaysia and Indonesia. Our clients are empowered with quality, cost-effective, and impactful end-to-end application development, like mobile and web applications, and cloud technology that remove technology roadblocks and increase their business efficiency. We believe in positively impacting the lives of people around us and the environment we live in through the use of technology. Hence, we are committed to providing a conducive environment for all employees to realise their full potential, who in turn have the opportunity to continuously drive innovation. We are searching for our next team members to join our growing team. If you love the idea of being part of a growing company with exciting prospects in mobile and web technologies that create positive impact on people’s lives, then we would love to hear from you. Co-Development Business Unit is looking for Data Engineer (Observability Engineer) This is a 1 - year contract role. Internal Code: A26323 Digital Excellence & Products Division (DXD) is a GovTech team within the Ministry of Education (MOE). DXD sits at the intersection of technology, design, and education, building meaningful products, platforms, and digital services that improve teaching, learning, school operations, and the experience of students, teachers, and school leaders. What will you do? We are looking for an Observability Engineer to help build and operate the observability capabilities of the future SSOE platform. You will help provide end-to-end visibility across MOE's technology environment, spanning on-premises infrastructure, networks, applications, cloud platforms, and hybrid environments. You will enable engineering and operations teams to understand system health, identify issues early, diagnose incidents quickly, and continuously improve service reliability. As an Observability Engineer, you will establish and operate consistent observability capabilities across SSOE infrastructure and applications. You will work across metrics, events, logs, and traces to provide a unified view of service health and performance. You will define instrumentation standards, service-level indicators and objectives, alerting strategies, dashboards, and operational health signals. You will work closely with application, infrastructure, network, security, and platform teams to ensure observability is built into services from the outset rather than added after deployment. Observability Engineering Design and operate end-to-end observability across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments Collect, aggregate, and correlate metrics, events, logs, and traces across infrastructure and application workloads Define and maintain observability standards that work consistently across legacy, on-premise, containerised, and cloud-native systems Establish application and infrastructure instrumentation standards using OpenTelemetry and other appropriate technologies Support engineering teams with instrumentation, SDK, agent, and telemetry integration Define common conventions for service naming, metadata, tagging, correlation IDs, and telemetry enrichment Identify observability gaps and continuously improve end-to-end visibility across SSOE services Service Reliability & Monitoring Define SLIs, SLOs, alerting rules, and service health indicators for critical services Build operational dashboards covering infrastructure health, application performance, user experience, availability, capacity, and service reliability Develop leadership-level views that provide meaningful visibility into service performance and operational trends Design actionable alerting that enables teams to identify and respond to issues while minimising unnecessary alert noise Establish monitoring and operational-readiness requirements for new applications, infrastructure, and platform components Use observability data to support capacity planning, performance analysis, reliability improvements, and operational decision-making Telemetry & Integration Define secure telemetry collection and routing across on-premise environments, GCC, cloud platforms, and approved SaaS services Work with infrastructure and platform teams to integrate telemetry from servers, network devices, applications, containers, databases, and managed cloud services Design observability approaches that account for network boundaries, security zones, data residency, and connectivity constraints Define telemetry retention, lifecycle, and cost-management requirements Ensure logs, metrics, and traces can be correlated across distributed and hybrid systems Incident Management & Continuous Improvement Support operational teams during incidents by using observability data to identify symptoms, dependencies, and potential root causes Participate in incident investigation, root-cause analysis, and post-incident reviews Identify recurring operational issues and recommend improvements to instrumentation, alerting, architecture, or operational processes Define appropriate SLOs and operational health indicators for observability services Participate in operational support and on-call responsibilities for owned services Maintain architecture documentation, operational procedures, and runbooks What are we looking for? Minimum 3–5 years of experience in observability engineering, Site Reliability Engineering (SRE), platform engineering, infrastructure engineering, or a related discipline At least 2 years of hands-on experience implementing or operating observability and monitoring capabilities in production environments Demonstrated experience working with metrics, logging, tracing, dashboards, alerting, and incident troubleshooting Experience monitoring and supporting production infrastructure, applications, or distributed systems Experience working with on-premise and/or cloud environments, with a
Employer contacts (email/phone/telegram) are hidden from the public preview — send your CV, and we will connect you directly.
Urgent question? Message @termdocs