SRE Lead (NJ) at Diverse Lynx in Jersey City, NJ
- Company: Diverse Lynx
- Location: Jersey City, NJ
- Posted: Sep 18, 2026
- Type: Full-time
- Experience: 10+ years
Overview
We are seeking an experienced SRE Lead to join our team in Jersey City, NJ. The ideal candidate will have strong experience in Site Reliability Engineering, cloud infrastructure, DevOps, automation, observability, and production support. This role requires a hands-on technical leader who can drive r…
Job description
- We are seeking an experienced SRE Lead to join our team in Jersey City, NJ. The ideal candidate will have strong experience in Site Reliability Engineering, cloud infrastructure, DevOps, automation, observability, and production support. This role requires a hands-on technical leader who can drive reliability, scalability, performance, and availability of enterprise applications.
Responsibilities
- Lead Site Reliability Engineering initiatives focused on application availability, scalability, performance, and resiliency.
- Design, implement, and maintain highly reliable and scalable production environments.
- Develop and maintain automation for infrastructure, deployments, monitoring, and operational processes.
- Establish and improve SLOs, SLIs, and SLAs and drive reliability metrics across applications and services.
- Lead incident management, troubleshooting, root-cause analysis, and post-incident reviews.
- Implement and maintain comprehensive monitoring, logging, alerting, and observability solutions.
- Work closely with development, QA, DevOps, infrastructure, and security teams to improve system reliability.
- Automate repetitive operational tasks and reduce manual intervention through scripting and infrastructure automation.
- Support CI/CD pipelines and implement reliability and quality checks throughout the software delivery lifecycle.
- Analyze system performance, identify bottlenecks, and implement solutions to improve application performance and stability.
- Participate in production deployments, release management, capacity planning, and disaster-recovery initiatives.
- Mentor engineers and provide technical leadership on SRE best practices.
- Define and document operational procedures, runbooks, troubleshooting guides, and disaster-recovery processes.
Requirements
- 10+ years of experience in SRE, DevOps, Production Engineering, Infrastructure Engineering, or a related field.
- Strong experience with Site Reliability Engineering principles and practices.
- Hands-on experience with AWS, Azure, or Eligible to workP cloud platforms.
- Strong knowledge of Kubernetes and Docker.
- Experience with Terraform or other Infrastructure as Code tools.
- Strong scripting/programming experience with Python, Shell scripting, or similar languages.
- Hands-on experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Strong experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, Datadog, New Relic, or ELK.
- Experience with production incident management, troubleshooting, and root-cause analysis.
- Strong understanding of Linux/Unix systems, networking, APIs, and distributed systems.
- Experience with Git and modern software development practices.
- Strong understanding of SLOs, SLIs, SLAs, error budgets, and reliability engineering.
- Experience working in Agile/Scrum environments.
- Excellent communication, problem-solving, and leadership skills.
Skills
Required
- Kubernetes and Docker
- Agile/Scrum environments