Description
Responsibilities Linux Systems & Automation (Core) - Manage large-scale Linux environments: troubleshooting and root-cause analysis - Write maintainable, hand-off-ready Bash / Ansible / Python automation - On-call for infrastructure, CI/CD, and production service incidents HPC Cluster & Storage - Operate HPC clusters (Slurm) along with usage analytics, auditing, and monitoring tools - Maintain and plan storage for compute environments (Lustre, NAS) Cloud & Hybrid Infrastructure - Manage multi-cloud environments (AWS, Alibaba Cloud, GCP) with Terraform / AWS CDK - Build and operate Docker (ECS) / Kubernetes (EKS) environments and their deployment workflows CI/CD & Developer Experience - Operate self-hosted GitLab server and Runner fleet - Operate CI/CD systems and design deployment pipelines for research and other projects GenAI / Internal Platform - Build internal AI platforms (LangChain / LangGraph / Bedrock, Elasticsearch RAG) - Develop MCP servers, chatbots, AI agents, and similar services
Requirements
- **5+ years** of hands-on Linux systems administration and infrastructure operations experience - Solid Linux internals knowledge (process / memory / filesystem / networking / systemd / cgroup); able to localize issues even without complete logs - Strong Bash / Shell scripting skills — able to write maintainable scripts that others can pick up - Programming ability for data processing, CLI tools, and API services; Python proficiency preferred - Solid storage fundame
Employer contacts (email/phone/telegram) are hidden from the public preview —
send your CV, and we will connect you directly.