Lead Site Reliability Engineer at SimCorp in Hyderabad
- Company: SimCorp
- Location: Hyderabad
- Posted: Sep 24, 2026
- Type: Full-time
- Experience: 6+ years
Overview
Join some of the most innovative thinkers in FinTech as we lead the evolution of financial technology. If you are an innovative, curious, collaborative person who embraces challenges and wants to grow, learn, and pursue outcomes with our prestigious financial clients, say Hello to SimCorp!
Job description
- Join some of the most innovative thinkers in FinTech as we lead the evolution of financial technology. If you are an innovative, curious, collaborative person who embraces challenges and wants to grow, learn, and pursue outcomes with our prestigious financial clients, say Hello to SimCorp!
- At its foundation, SimCorp is guided by our values — caring, customer success-driven, collaborative, curious, and courageous. Our people-centered organization focuses on skills development, relationship building, and client success. We take pride in cultivating an environment where all team members can grow, feel heard, valued, and empowered.
- If you like what we’re saying, keep reading!
- As a Lead Site Reliability Engineer, you will join one of our Product Areas and become part of a collaborative team focused on operating and improving our Azure-based SaaS platform. You’ll work with experienced engineers and cross-functional colleagues to help ensure reliability, scalability, and operational excellence for both onboarding and running client environments.
- This role is a launchpad for your development as an SRE—combining hands-on engineering, automation, monitoring, and platform support—with strong mentorship and continuous learning opportunities built in.
Responsibilities
- • Work hands-on across platform and infrastructure SRE activities — cloud infrastructure, Kubernetes/container platforms, CI/CD, and production systems — not just reviewing or delegating this work, but building and operating it directly.
- • Define, own, and report on SLIs, SLOs, and SLAs for critical services; manage error budgets and use them to drive concrete engineering and release decisions.
- • Actively participate in cost optimization efforts across cloud and platform infrastructure, identifying and implementing efficiency and rightsizing opportunities (FinOps practices).
- • Lead the design and implementation of strategies to ensure the reliability, scalability, and performance of critical systems and services.
- • Manage incident response, troubleshooting, and root cause analysis for system outages and performance issues, ensuring timely resolution and prevention of future incidents.
- • Develop and maintain monitoring, alerting, and automation systems to enhance service reliability and reduce manual intervention, including automation and AI-driven engineering approaches (e.g. AI-assisted anomaly detection, automated remediation, AI-supported incident triage) to reduce toil and speed up resolution.
- • Collaborate with development, operations, and product teams to optimize application performance and system infrastructure.
- • Drive initiatives for capacity planning, resource management, and scalability to meet growing business needs.
- • Implement continuous improvement processes to enhance the reliability and efficiency of systems and services.
- • Mentor and guide junior site reliability engineers, fostering a culture of collaboration and knowledge-sharing, while remaining personally hands-on with the underlying systems.
- • Participate in on-call rotations, providing leadership during critical incidents and ensuring minimal downtime.
- • Create and maintain documentation for incident management, system configurations, and operational processes.
- • Drive the adoption of industry best practices, tools, and technologies — including automation and AI-driven engineering tooling — to enhance system reliability and operational performance.
Requirements
- Lead Site Reliability Engineer role
- Hands-on experience across platform and infrastructure SRE activities — cloud infrastructure, Kubernetes/container platforms, CI/CD, and production systems
- Experience defining, owning, and reporting on SLIs, SLOs, and SLAs for critical services; managing error budgets
- Experience in cost optimization efforts across cloud and platform infrastructure (FinOps practices)
- Experience leading design and implementation of strategies for reliability, scalability, and performance
- Experience managing incident response, troubleshooting, and root cause analysis
- Experience developing and maintaining monitoring, alerting, and automation systems
- Experience collaborating with development, operations, and product teams
- Experience driving initiatives for capacity planning, resource management, and scalability
- Experience implementing continuous improvement processes
- Experience mentoring and guiding junior site reliability engineers
- Experience participating in on-call rotations
- Experience creating and maintaining documentation for incident management, system configurations, and operational processes
- Experience driving adoption of industry best practices, tools, and technologies
Skills
Required
- Site reliability engineering
- Cloud infrastructure (Azure)
- Kubernetes/container platforms
- CI/CD
- Production systems
- SLIs, SLOs, SLAs
- Error budgets
- Cost optimization (FinOps)
- Incident response
- Troubleshooting
- Root cause analysis
- AI-assisted anomaly detection
Benefits
- Attractive salary, bonus scheme, and pension are essential for any work agreement. However, in SimCorp we believe we can offer more. Therefore, in addition to the traditional benefit scheme, we provide a good work-life balance: flexible working hours and a hybrid model. Simcorp follows a global hybrid policy, asking employees to work from the office two days each week while allowing remote work on other days.
- Simcorp does offer opportunities for professional development: there is never just only one route - we offer an individual approach to professional development to support the direction you want to take.
About SimCorp
For over 50 years, we have worked closely with investment and asset managers to become the world’s leading provider of integrated investment management solutions. We are 3,000+ colleagues with a broad range of nationalities, educations, professional experiences, ages, and backgrounds. SimCorp is an independent subsidiary of the Deutsche Börse Group. Following the recent merger with Axioma, we leverage the combined strength of our brands to provide an industry-leading, full, front-to-back offering for our clients.