Описание
Job Summary We are seeking an experienced Site Reliability Engineer (SRE) / Production Support Engineer with 10+ years of experience in enterprise infrastructure, production support, system administration, and SRE operations. The ideal candidate will have strong hands-on experience in AIX, Linux/RHEL, application production support, monitoring, automation, incident management, and cloud/DevOps environments. The candidate will be responsible for ensuring high availability, reliability, performance, and security of critical enterprise applications and infrastructure in a 24x7 production support environment.
The role
requires close collaboration with development, infrastructure, middleware, security, and business teams to identify and resolve production issues and continuously improve service reliability. Key
Responsibilities
Site Reliability Engineering Define and establish SLIs, SLOs, SLAs, error budgets, MTTD, and MTTR for enterprise applications and services. Monitor application availability, latency, performance, capacity, and overall reliability. Identify and implement Golden Signals and observability best practices across applications and infrastructure.
Analyze production incidents and recurring issues to identify root causes and implement permanent remediation. Develop automation and engineering solutions to reduce manual operational activities and improve reliability. Participate in 24x7x365 production support and on-call operations.
Perform capacity planning and proactively identify infrastructure and application resource requirements. Support highly available and resilient application architectures. Participate in disaster recovery planning, testing, and implementation.
Production & Infrastructure Support Provide L2/L3 production and infrastructure support for critical enterprise applications. Perform administration, troubleshooting, configuration, maintenance, and performance tuning of IBM AIX and RHEL/Linux servers. Support server patching, upgrades, maintenance, and DR activities.
Troubleshoot application, middleware, operating system, connectivity, and infrastructure-related issues. Perform middleware administration and restart activities for WebSphere, JBoss, IBM HTTP Server, and Apache Tomcat. Support firewall changes, SSL certificate renewals, SCP configuration, and system-to-system key exchanges.
Support application deployment activities across Blue/Green environments. Coordinate application-related changes and deployments across development, infrastructure, middleware, and business teams. Monitoring & Observability Configure and maintain monitoring and alerting for infrastructure and application services.
Work with monitoring and observability platforms such as: Grafana Centreon AppDynamics ELK / Kibana Application and infrastructure logging platforms Analyze system and application metrics, logs, and performance trends. Develop appropriate alerts to proactively identify service degradation and failures. Promote observability practices and help development teams implement effective monitoring.
Incident & Change Management Manage and resolve incidents within defined SLA/OLA timelines. Participate in major incident and emergency response activities. Perform incident investigation, troubleshooting, root-cause analysis, and problem management.
Raise and manage Change Requests (CRs) for application deployments, BAU fixes, infrastructure changes, middleware changes, and maintenance activities. Create and manage service requests and incident tickets. Prepare RCA and corrective/preventive action plans for recurring and high-priority incidents.
Follow ITIL-based Incident, Change, Problem, and Service Request Management processes. Automation & DevOps Develop automation scripts using Python and Shell scripting to improve operational efficiency. Automate repetitive infrastructure and production support activities.
Work with DevOps tools and technologies including: Git / GitHub / Bitbucket Jenkins Ansible Chef Docker Maven JFrog / Nexus Repository SonarQube Fortify / Nexus IQ Support CI/CD pipelines and application deployment processes. Collaborate with development teams to integrate monitoring, logging, and reliability controls into deployment pipelines. Cloud & Application Support Provide support for applications hosted on AWS and Pivotal Cloud Foundry (PCF) environments.
Work with enterprise application technologies including: IBM WebSphere IBM HTTP Server JBoss Apache Tomcat Java applications Support microservices-based applications and their associated infrastructure. Assist with application migration, optimization, and adoption of new technologies where required. Security & Compliance Apply system security best practices across AIX and Linux environments.
Support security hardening, patching, access management, and vulnerability remediation. Work with security and access management tools such as IBM Tivoli Access Manager. Support SSL certificate management and renewal activities.
Prepare quarterly operational reports and documentation required for audit, risk, and compliance activities. Ensure production changes and operational activities comply with organizational security and governance standards. x Red Hat Enterprise Linux (RHEL) Linux Windows Server SRE / Monitoring SLI / SLO / SLA Error Budgets MTTD / MTTR Observability Golden Signals Grafana Centreon AppDynamics ELK / Kibana Cloud & Containers AWS Pivotal Cloud Foundry (PCF) Docker Middleware & Application Technologies IBM WebSphere Application Server IBM HTTP Server JBoss Apache Tomcat Java Databases DB2 UDB SQL MySQL / MariaDB Automation & Scripting Python Shell Scripting Groovy DevOps / CI-CD Git / GitHub Bitbucket Jenkins Ansible Chef Maven JFrog / Nexus Repository SonarQube Fortify Nexus IQ ITSM / Ticketing ServiceNow BMC Remedy IBM ISM JIRA Soft Skills Strong analytical and troubleshooting skills.
Excellent incident management and problem-solving capa
Контакты работодателя (email/phone/telegram) скрыты из публичного превью —
отправьте резюме, чтобы мы связали вас напрямую.