AWS Platform Operations Lead, Vice President at State Street in Quincy, Massachusetts
- Company: State Street
- Location: Quincy, Massachusetts
- Posted: Sep 19, 2026
- Type: Full-time
- Salary: $120,000 - $217,500 Annual
- Experience: 12+ years
- Visa sponsorship available
Overview
State Street is seeking a visionary AWS Federated Platform Operations Lead to lead the evolution of cloud operations toward an intelligent, automated, and AI-enabled operating model. This role is responsible for the operational strategy, reliability, resiliency, observability, security operations go…
Job description
- State Street is seeking a visionary AWS Federated Platform Operations Lead to lead the evolution of cloud operations toward an intelligent, automated, and AI-enabled operating model.
- This role is responsible for the operational strategy, reliability, resiliency, observability, security operations governance, and operational transformation of the AWS Federated platform. The successful candidate will partner closely with Platform Engineering, Architecture, Security, AI, and Enterprise Operations teams to deliver a secure, resilient, scalable, and highly automated cloud platform supporting the firm's most critical business workloads.
- This is not a traditional infrastructure operations leadership role. Instead, the AWS Federated Platform Operations Lead will drive the next generation of cloud operations by leveraging AI, agentic workflows, automation, operational intelligence, reliability engineering, and platform-based support models to improve service quality, reduce operational risk, and accelerate innovation.
- The role serves as a key member of the AWS Federated leadership team and will play a pivotal role in shaping the future multi-cloud operating model across State Street.
Responsibilities
- Define and Execute the Cloud Operations Strategy
- Develop and execute the long-term operational vision and roadmap for AWS Federated.
- Establish operational standards, governance frameworks, and service management practices aligned with enterprise objectives.
- Drive continuous improvement in operational maturity, platform reliability, customer experience, and service quality.
- Partner with engineering leaders to integrate operability, observability, resiliency, and supportability into platform design decisions.
- Influence enterprise cloud operating model strategy and operational transformation initiatives.
- Lead the Transformation to AI-Powered Operations
- Drive adoption of Agentic AI capabilities across cloud operations and platform management.
- Establish a roadmap for autonomous cloud operations utilizing AI agents, automation, and operational intelligence platforms.
- Incident correlation
- Root cause analysis
- Capacity forecasting
- Operational insights
- Risk identification
- Automated remediation
- Develop AI-powered operational assistants that improve engineering productivity and accelerate troubleshooting.
- Partner with AI platform teams to operationalize emerging capabilities including foundation models, operational copilots, workflow orchestration, and intelligent automation.
- Establish governance, controls, and guardrails for the responsible adoption of AI within operational processes.
- Drive Reliability Engineering and Platform Resilience
- Establish and mature Site Reliability Engineering (SRE) practices across AWS Federated.
- Define service-level objectives (SLOs), error budgets, reliability scorecards, and resilience metrics.
- Build platform reliability programs focused on prevention rather than reaction.
- Disaster recovery
- Recovery automation
- Chaos engineering
- Failure testing
- Cyber resiliency validation
- Ensure operational readiness requirements are consistently incorporated into platform architecture and service enablement initiatives.
- Define Next-Generation Observability and Operational Intelligence
- Establish the strategic direction for observability across cloud, infrastructure, application, data, and security domains.
- Create a unified operational intelligence framework that combines telemetry, operational events, security signals, and platform insights.
- Distributed tracing
- Metrics aggregation
- Log analytics
- AI-driven anomaly detection
- Predictive analytics
- Leverage machine learning and analytics to proactively identify capacity, performance, reliability, and operational risks.
- Enable data-driven operational decision making through real-time dashboards and executive reporting.
- Accelerate Automation and Platform Productivity
- Establish Operations-as-Code and Automation-as-a-Product principles across the platform.
- Lead initiatives to eliminate repetitive operational activities through automation and self-service capabilities.
- Develop reusable automation frameworks that standardize platform operations and reduce operational complexity.
- Drive adoption of event-driven automation and intelligent workflows.
- Measure and improve operational efficiency through AI-assisted engineering, automation, and workflow optimization.
- Strengthen Security, Compliance, and Operational Risk Management
- Partner with Cyber Security, Risk, and Compliance teams to continuously improve platform security posture.
- Drive operational governance for vulnerability remediation, patch strategy, cloud controls, and regulatory compliance requirements.
- Implement automated compliance and continuous assurance capabilities wherever feasible.
- Support cyber resilience and cyber-immunity initiatives through automation, monitoring, and operational controls.
- Ensure cloud operational practices remain aligned with evolving enterprise security standards.
- Provide Strategic Leadership During Critical Operational Events
- Serve as the senior operational leader for AWS Federated during significant service-impacting events.
- Provide strategic decision-making and executive communication during major incidents and platform disruptions.
- Drive post-event learning, systemic improvement initiatives, and reliability investments.
- Partner with engineering teams to ensure long-term corrective actions are implemented and measured.
- Leadership and Organizational Development
- Build and lead a high-performing organization spanning reliability engineering, observability, operational intelligence, automation, and platform operations disciplines.
- Foster a culture of innovation, accountability, engineering excellence, and continuous improvement.
- Develop future leaders with expertise in cloud engineering, AI-enabled operations, resiliency engineering, and platform management.
- Champion adoption of emerging technologies and modern engineering practices across the organization.
Requirements
- Required Experience
- Bachelors degree
- 12+ years of experience in cloud platform engineering, reliability engineering, operations transformation, infrastructure engineering, or related disciplines.
- 5+ years of leadership experience managing large-scale technology organizations.
- Deep expertise operating enterprise-scale AWS environments in regulated industries.
- Proven experience establishing operational strategy for complex cloud platforms.
- Strong knowledge of cloud architecture, networking, security, identity, automation, and service management disciplines.
- Experience leading large-scale operational transformation initiatives.
- Demonstrated ability to influence senior executives and drive cross-functional change.
- Financial services or highly regulated industry experience.
- Experience implementing AIOps, operational analytics, or observability platforms.
- Experience with platform engineering and internal developer platforms.
- Familiarity with Agentic AI architectures, AI orchestration platforms, and intelligent automation technologies.
- Experience with Terraform, Infrastructure as Code, GitOps, and modern CI/CD practices.
- AWS Professional or Specialty Certifications.
- Experience leveraging generative AI to improve engineering productivity and operational outcomes.
Skills
Required
- Cloud platform engineering
- Reliability engineering
- Operations transformation
- Infrastructure engineering
- AWS environments
- Cloud architecture
- Networking
- Security
- Identity
- Automation
- Service management
- Operational strategy
Preferred
- AIOps
- Operational analytics
- Observability platforms
- Platform engineering
- Internal developer platforms
- Agentic AI architectures
- AI orchestration platforms
- Intelligent automation technologies
Benefits
- $120,000 - $217,500 Annual
- The range quoted above applies to the role in the primary location specified. If the candidate would ultimately work outside of the primary location above, the applicable range could differ.
- Employees are eligible to participate in State Street’s comprehensive benefits program, which includes: our retirement savings plan (401K) with company match; insurance coverage including basic life, medical, dental, vision, long-term disability, and other optional additional coverages; paid-time off including vacation, sick leave, short term disability, and family care responsibilities; access to our Employee Assistance Program; incentive compensation including eligibility for annual performance-based awards (excluding certain sales roles subject to sales incentive plans); and, eligibility for certain tax advantaged savings plans.
- For a full overview, visit https://hrportal.ehr.com/statestreet/Home.
About State Street
Across the globe, institutional investors rely on us to help them manage risk, respond to challenges, and drive performance and profitability. We keep our clients at the heart of everything we do, and smart, engaged employees are essential to our continued success. We are committed to fostering an environment where every employee feels valued and empowered to reach their full potential. As an essential partner in our shared success, you’ll benefit from inclusive development opportunities, flexible work-life support, paid volunteer days, and vibrant employee networks that keep you connected to what matters most. Join us in shaping the future.