Site Reliability Engineer (Azure Preferred) at Fidelity National Information Services in IND PUNE FL7
- Company: Fidelity National Information Services
- Location: IND PUNE FL7
- Posted: Sep 17, 2026
- Type: Full-time
- Experience: 5+ years
Overview
As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, performance, and availability of mission-critical banking, payments, and capital markets platforms. You will drive automation, strengthen operational resilience, and improve observability across cloud-…
Job description
- As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, performance, and availability of mission-critical banking, payments, and capital markets platforms. You will drive automation, strengthen operational resilience, and improve observability across cloud-native and distributed systems. Working closely with engineering, DevOps, security, QA, and product teams, you will help deliver highly available services while reducing operational risk. Success in this role is measured through platform stability, service reliability, incident reduction, and continuous operational improvement.
Responsibilities
- Design, implement, and maintain monitoring and observability solutions for infrastructure, applications, and customer experience.
- Build and enhance automation frameworks to improve operational efficiency and reduce manual processes.
- Ensure high availability, reliability, scalability, and performance of critical production systems.
- Lead incident management activities, including triage, root cause analysis, recovery, and post-incident reviews.
- Perform capacity planning, performance tuning, and infrastructure optimization to support business growth.
- Develop and manage Infrastructure as Code solutions for consistent and scalable cloud deployments.
- Maintain and optimize CI/CD pipelines to enable reliable and secure software delivery.
- Collaborate with security teams to implement platform security controls and compliance best practices.
- Develop, validate, and improve disaster recovery, backup, and business continuity strategies.
- Partner with engineering, DevOps, QA, and product teams to achieve service-level objectives and operational excellence.
- Participate in on-call rotations and provide support for critical production environments.
Requirements
- 5+ years of experience in Site Reliability Engineering, Production Support, Platform Engineering, DevOps, Cloud Operations, or a related field.
- Hands-on experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform.
- Strong knowledge of Infrastructure as Code and automation tools such as Terraform and Ansible.
- Experience supporting web applications, APIs, distributed systems, and modern software architectures.
- Proficiency with monitoring and observability tools including Prometheus, Grafana, Datadog, or similar platforms.
- Experience with logging and analytics solutions such as Splunk, ELK Stack, or equivalent technologies.
- Strong scripting and automation skills using Python, Bash, or similar programming languages.
- Experience designing, maintaining, and optimizing CI/CD pipelines using Jenkins, GitLab CI/CD, Azure DevOps, or related tools.
- Knowledge of containerization technologies such as Docker and container orchestration platforms.
- Demonstrated experience in incident management, root cause analysis, and production support in enterprise environments.
- Experience in applying SRE practices, reliability engineering principles, and service-level management in large-scale environments.
- Knowledge of Kubernetes, cloud-native architectures, and microservices-based platforms.
- Experience conducting operational readiness assessments and post-mortem reviews.
- Understanding of disaster recovery, high-availability design patterns, and resiliency engineering.
- Industry certifications in AWS, Azure, Google Cloud, Kubernetes, DevOps, or Site Reliability Engineering
Skills
Required
- Site Reliability Engineering
- Production Support
- Platform Engineering
- DevOps
- Cloud Operations
- Logging and analytics: Splunk, ELK Stack
- Scripting and automation: Python, Bash
- Containerization: Docker
- Container orchestration
- Incident management
- Root cause analysis
- Production support
Preferred
- SRE practices
- Reliability engineering principles
- Service-level management
- Kubernetes
- Cloud-native architectures
- Microservices-based platforms
- Operational readiness assessments
- Post-mortem reviews
Benefits
- A work environment built on collaboration, flexibility and respect
- Competitive salary and attractive range of benefits designed to help support your lifestyle and wellbeing
- Varied and challenging work to help you grow your technical skillset