AI/MLOps SRE Lead Engineer at Regeneron Pharmaceuticals in Hyderabad
- Company: Regeneron Pharmaceuticals
- Location: Hyderabad
- Posted: Sep 23, 2026
- Type: Full-time
- Experience: 8+ years
Overview
At Regeneron Digital & Technology, we are expanding our AI and Platform Engineering capabilities to support next-generation intelligent systems, machine learning platforms, and cloud-native technologies. We are seeking an AI/MLOps SRE Lead Engineer to drive reliability, scalability, observability, a…
Job description
- At Regeneron Digital & Technology, we are expanding our AI and Platform Engineering capabilities to support next-generation intelligent systems, machine learning platforms, and cloud-native technologies. We are seeking an AI/MLOps SRE Lead Engineer to drive reliability, scalability, observability, and operational excellence across our AI, ML, and cloud ecosystem. This role will lead the design and operation of resilient platforms supporting machine learning workloads, LLMs, AI Agents, and enterprise-scale automation while enabling engineering teams to innovate with speed and confidence.
Responsibilities
- Discover your role
- Drive service reliability, availability, and performance across multi-cloud environments by establishing SLOs, SLIs, error budgets, and reliability standard methodologies.
- Design, build, and operate enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
- Develop AI-powered observability capabilities using anomaly detection, predictive analytics, and automated remediation to proactively identify and resolve operational issues.
- Lead the implementation, monitoring, and optimization of LLM, SLM, RAG, and AI Agent platforms, ensuring performance, governance, scalability, and operational excellence.
- Design and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation solutions to improve engineering productivity and platform resilience.
- Architect enterprise ChatOps solutions integrating operational events, observability platforms, AI workflows, and automated remediation capabilities.
- Partner closely with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML solutions.
- Evaluate emerging AI-native operational technologies and integrate innovative solutions that enhance platform reliability, engineering efficiency, and business value.
- Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to improve enterprise platform maturity.
- Serve as a technical leader and trusted advisor, mentoring engineers and influencing reliability engineering, MLOps, cloud platform strategy, and AI-enabled operations across the organization.
Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related field; Master's degree preferred.
- 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology subject areas with enterprise-scale delivery experience.
- Strong hands-on experience operating across two or more major cloud platforms, including AWS, GCP, and Azure.
- Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
- Proven experience implementing end-to-end ML workflows including model training, experiment tracking, deployment, monitoring, and pipeline orchestration.
- Hands-on experience applying machine learning techniques such as anomaly detection, predictive analytics, time-series modelling, and operational intelligence within enterprise environments.
- Advanced proficiency with Infrastructure as Code tools such as Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
- Strong programming and scripting skills in Python, Go, Bash, or similar languages.
- Experience designing enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, logging, and monitoring platforms.
- Demonstrated expertise in anomaly detection, predictive analytics, automated remediation, and AI-assisted operational capabilities.
- Proven experience designing and implementing enterprise ChatOps solutions, including operational workflow automation and AI-enabled integrations.
- Strong ability to identify technical debt, assess platform risks, influence technical strategy, and drive modernization initiatives.
- Experience using AI tools, AI Agents, and LLM-powered assistants to improve engineering operations, incident management, and developer productivity.
- Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS preferred.
- Familiarity with MLOps technologies including Kubeflow, Feast, MLflow, and model evaluation frameworks such as LangSmith, RAGAS, Evidently AI, or Weights & Biases preferred.
- Knowledge of cloud cost optimization, policy-as-code, compliance automation, FinOps practices, and multi-cloud governance preferred.
Skills
Required
- Site Reliability Engineering
- Platform Engineering
- DevOps
- Cloud platforms (AWS, GCP, Azure)
- CI/CD automation
- Programming languages (Python, Go, Bash)
- ChatOps solutions
- Technical debt assessment
Benefits
- Regeneron offers a competitive and comprehensive total rewards package which may include, depending on country and role: annual bonuses or other incentive plans, equity awards, pension or retirement benefits, 401(k) company match, health and wellness programs, fitness centers, insurance benefits (e.g. medical, dental, vision, life and disability), paid time off, and family support benefits.
- For additional information about Regeneron benefits in the U.S., please visit https://careers.regeneron.com/en/working-at-regeneron/total-rewards/. For other locations, additional information will be provided during the recruitment process. If you have any questions, please speak with your recruiter.
- Where necessary, we disclose salary ranges for roles in all countries in which we operate. The final offer will be determined within the relevant range based on the country of employment, specific role level, and your skills and experience.
- In some countries, collective bargaining agreements (CBAs) may apply and influence certain elements of pay or benefits.
About Regeneron Pharmaceuticals
Build our future together Regeneron is founded on the belief that the right idea, combined with the right team, can lead to significant transformations. Our growing global network is dedicated to inventing, developing, and commercializing medicines that change lives for those with serious diseases. In doing so, we are pioneering innovative approaches to science, manufacturing, and commercialization, as well as redefining our understanding of health.