Zorky CRMZorky CRM
EN|RU
@termdocs
← Все вакансии

Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

San Diego, USСкор 52/100сегодня
Аналитика рынка
📊 Data Engineer: зарплаты и спрос на рынке
Стек
kubernetesservicenow
Откликнуться
Загрузите резюме — мы свяжем вас с работодателем напрямую через нашу базу.
Отправить резюме →
Описание
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.Join us to put AI to work for people.  Position Location: This is a Flexible (Hybrid) position.  Flexible positions require 2 days per week in a ServiceNow office location.  We have offices in several locations, including San Francisco, CA; Pleasanton, CA; Santa Clara, CA; San Diego, CA; and Kirkland, WAPlease Note:  This position will include supporting our US Regulated Markets. “This position requires passing a ServiceNow background screening, USFedPASS (US Federal Personnel Authorization Screening Standards). This includes a credit check, criminal/misdemeanor check and taking a drug test. Any employment is contingent upon passing the screening.  Due to Federal requirements, only US citizens, US naturalized citizens or US Permanent Residents, holding a green card, will be considered. About the team: We are seeking a Senior Staff Software Engineer (IC5 Level) to join our Data Platform Engineering organization.   You will be a hands-on technical lead who will lead the design and delivery of shared, multi-tenant platform services — Postgres, queueing/streaming, and key-value stores — on Kubernetes, and serves as the technical anchor for their availability, resilience, and operating model.What you get to do in this role:You will lead the design and delivery of shared platform services — managed Postgres, queueing/streaming, and key-value/cache — that run on Kubernetes across a large global fleet and that product teams across the company depend on, owning them from design through production operation.You will act as the technical lead on major initiatives within your team, breaking down ambiguous problems — “offer HA Postgres as a service to every cluster,” “make our queueing tier survive a zone loss” — into clear, executable designs with explicit availability, durability, and cost targets.You will define the HA, failover, disaster-recovery, and multi-region architecture for the services you own, and the Kubernetes-native automation (operators, controllers, CRDs, self-service APIs) that provisions, scales, upgrades, and fails them over without human intervention.You will partner with principal and distinguished engineers to align your work with the broader platform architecture and standards, and with product teams to set consumption contracts, tenancy models, and SLOs for shared services.You will identify technical risks early and drive them to resolution, with a strong focus on reliability, scalability, and operability — leading failure-mode analysis, game days, and post-incident reviews for stateful systems.You will spend significant time hands-on — designing, coding, and reviewing the core systems your team builds, such as operators, controllers, infrastructure automation, and platform services.You will mentor mid-level and junior engineers and raise the engineering bar through code reviews, design feedback, and pairing — particularly around distributed-systems and data-service design.  To be successful in this role you have:Experience leveraging or critically thinking about how to integrate AI into engineering and platform work — whether using AI-powered tooling, automating operational workflows, building agentic systems for fleet visibility and operations, or reasoning about AI’s impact on how infrastructure is built and run.12+ years of software development experience with a Bachelor's degree; OR 8+ years with a Master's degree; OR 5+ years with a PhD; OR equivalent work experience.8+ years building production software, with solid experience operating distributed systems and running stateful workloads on Kubernetes at scale.A track record of leading the design and delivery of at least one shared, multi-tenant infrastructure service used broadly by other teams — a relational database service (Postgres or similar), a message queue or streaming platform (Kafka, NATS, RabbitMQ, or similar), or a key-value/cache service (Redis/Valkey, etcd, or similar) — including its HA, failover, and operating model.Deep understanding of high availability and failure handling: replication topologies, leader election, quorum and consensus, split-brain avoidance, backup/restore and point-in-time recovery, and designing to explicit RPO/RTO and durability targets.Strong system design skills — you can reason rigorously about consistency models, partitioning and rebalancing, capacity planning, tenant isolation, and noisy-neighbor mitigation, and communicate the trade-offs clearly to engineers and stakeholders.Hands-on experience with at least one major hyperscaler (AWS, Azure, GCP), including its core compute, networking, storage, and IAM primitives.Strong working knowledge of containers and Kubernetes (including stateful primitives: StatefulSets, CSI/persistent storage, PDBs, topology spread), CI/CD and GitOps-based delivery, and infrastructure-as-code.Strong programming skills in Go (or strong systems-language skills with a willingness to work primarily in Go).It also helps if you have:Experience building or extending Kubernetes operators that manage stateful systems (e.g., CloudNativePG, Zalando/Crunchy Postgres operators, Strimzi, Redis/Valkey operators), and opinions on when to adopt versus build.Deep expertise in one of the target systems: Postgres internals (WAL, streaming/logical replication, vacuum, connect
Контакты работодателя (email/phone/telegram) скрыты из публичного превью — отправьте резюме, чтобы мы связали вас напрямую.
Срочный вопрос? Напишите @termdocs