Zorky CRMZorky CRM
EN|RU
@termdocs
← All jobs

Site Reliability Engineer( SRE)

hybrid7,000 SGDSingapore, SGScore 67.5/100today
Stack
MitigationBusiness OptimizationSystem AvailabilityBusiness SupportMonitoring SolutionsPlatform ManagementAutomation ToolsSystem MonitoringOperational Delivery TeamsLoad BalancingPractical trainingManaging Daily Operations
Apply
Upload your CV — we will connect you with the employer directly through our pool.
Send your CV →
Description
Job Responsibilities Ensure the stability, reliability, and high availability of the company’s overseas production environment, continuously improving system availability and service quality. Manage resource provisioning, capacity planning, monitoring, change management, incident response, and daily operations to maintain business continuity and stability. Review system architecture and technical solutions, identify potential risks, and implement mitigations to optimize system stability, performance, and resource efficiency. Participate in on-call rotations, responding promptly to production incidents to safeguard business operations. Build and enhance observability systems, including monitoring, logging, and distributed tracing, to improve monitoring capabilities and fault detection efficiency. Develop and optimize automation platforms and engineering efficiency tools to advance operational automation and team delivery effectiveness. Explore and promote the application of AI technologies in operations scenarios, leveraging AI tools to improve automation, fault analysis, knowledge management, and R&D efficiency. Collaborate closely with R&D, product, security, and infrastructure teams to drive stability initiatives, implement best practices, and support ongoing business development. Job Requirements Bachelor’s degree or above in Computer Science, Software Engineering, Information Technology, or a related field, with 3+ years of experience in Site Reliability Engineering (SRE), Computer Systems Administrator, Platform Engineer, or Cloud Engineer. , Python, Go, Java, or C++), with strong software development and automation skills. , Alibaba Cloud, Azure, AWS, GCP) is a plus. Solid understanding of Linux, computer networking, load balancing, distributed systems, and high-availability architectures. Ability to quickly diagnose issues, communicate across teams, and drive solutions—developing system optimization and stability plans aligned with business goals, including dependency management, traffic governance, and disaster recovery planning. , Prometheus, Grafana, ELK, or similar), along with scripting knowledge (Bash or Python) and familiarity with CI/CD concepts is preferred. Familiarity with AI tools and their applications in software development, automated operations, or R&D efficiency—understanding of AI Agents or AIOps technologies; practical experience is a plus. Strong communication, teamwork, and project management skills; ability to adapt to a fast-paced technical environment and continuously learn and apply new technologies.
Employer contacts (email/phone/telegram) are hidden from the public preview — send your CV, and we will connect you directly.
Urgent question? Message @termdocs