110 sre manager jobs at 66 companies in Paramus, NJ
1mo
Save
Mark Applied
Hide
1mo
SRE Manager, ML Operations
New York City, New York, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Manage site reliability and ML operations for Apple's advertising platform; role emphasizes building and operating reliable ad services across Apple products.
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's or military experience; 4+ years in technology management/software engineering/SRE/cyber risk; 2+ years with cloud (AWS/GCP/Azure); 1+ year with open-source programming; strong communication and analytical skills.
Akoya: Secure API network for sharing consumer financial data.
Deep AWS, Kubernetes/EKS, hybrid networking, CI/CD, observability, security, scripting (Python/Go/Bash), and incident leadership experience for large multi-account, multi-region environments.
Cary or New York City or Tampa or Bridgewater or Clarks Summit or Greenville
$120k-$170k/yrHybridFull Time
MetLifeNYSE: MET: Global provider of insurance, annuities, and financial services.
7+ YOE7+ years mainframe engineering experience with expertise in IBM z/OS, COBOL, CICS, z/OS Connect; proven leadership in mainframe modernization and cross-functional team delivery.
IBM z/OS, COBOL, CICS, z/OS Connect, zLinux, OpenShift, RedHat Ansible for IBM Z Collections, Ansible, Python, MQ
Kontakt.io: Provides AI-powered IoT platforms for hospital operational intelligence.
10+ YOE10+ years in SRE/cloud infrastructure, deep AWS/Kubernetes expertise, monitoring/observability (Prometheus/Grafana/OpenTelemetry/Datadog), Terraform/IaC, incident management, and leadership experience.
8+ YOE3+ MgmtBachelor's in CS or equivalent, 8+ years software development, 3+ years people management, experience with distributed systems, on-call and postmortem practices, strong communication and critical thinking.
Coralogix: AI-powered observability and security data platform.
5+ YOE2+ MgmtTeam lead with 5+ years SRE/DevOps experience, 2+ years lead experience; Kubernetes, AWS, Terraform/Crossplane, monitoring tools, FedRAMP familiarity preferred; strong leadership and incident management skills; EST/CT timezone.
Technical Product Manager II, Site Reliability Engineering
New York City, New York, United States
$120k-$142k/yrHybridFull Time
The New York TimesNYSE: NYT: Publishes global journalism, digital subscriptions, and lifestyle media products.
5+ YOE5+ years of product management in platform, infrastructure, SRE, or similar tech domains; knowledge of SRE practices; testing/observability in cloud-native environments; ability to turn data into roadmaps and impact.
Voya FinancialNYSE: VOYA: Provides retirement, investment, and insurance products and services.
Experience leading large-scale cloud migration or data center exit programs; strong technical understanding of infrastructure, security, networking, databases, and cloud migration patterns; ability to coordinate cross-functional teams and manage risks and dependencies.
Azure, AWS, Google Cloud Platform, DevSecOps, CI/CD, SRE
Basis: AI agents that automate complex accounting and tax workflows.
5+ YOE5+ years building and operating production infrastructure; strong software engineering; cloud, networking, databases, security; IaC, CI/CD, containerization; observability and incident management; on-call and incident leadership.
Optimal Market Technologies: Operates a broker-dealer platform for wholesale options execution.
Hands-on Linux and network administration, strong scripting/automation, Infrastructure-as-Code experience, production support and incident response ownership, staff management/mentoring, familiarity with C++, Python, SQL, Azure, and real-time trading systems.
C++, Python, SQL, Linux, CentOS7, RHEL9, Microsoft Azure, Microsoft Azure Virtual Desktop (AVD), Claude Code, FIX Protocol, PostgreSQL, AERON
Laravel: A PHP web framework for building modern web applications.
6+ YOE6+ years leading infrastructure engineering orgs (manager-of-managers), SRE practices, Kubernetes, AWS, Infrastructure as Code (Terraform/CDK), strong leadership and communication.
Charlotte or Tempe or New York City or Fort Mill or Austin or Boston or San Diego or Saint Louis
$73k-$171k/yrOnsiteFull Time
Perficient: Provides digital transformation and AI consulting for global enterprises.
8+ YOE5+ Mgmt8+ years in IT operations/SRE/infrastructure,5+ years leading incident management; bachelor's in CS/IT/engineering or equivalent; experience with SRE, ITSM/ITIL, ServiceNow, Dynatrace, cloud platforms, AIOps and automation.
Bellevue or Livingston or New York City or Sunnyvale or San Francisco
$182k-$242k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOE3+ Mgmt3+ years engineering management and 7+ years technical experience in cloud operations/SRE; strong knowledge of cloud platforms, K8S, observability, incident management, and systems programming (Go); experience defining SLAs/SLOs and running on-call.
CAVA GroupNYSE: CAVA: Operates a chain of Mediterranean fast-casual restaurants.
5+ YOE2+ MgmtLead cloud engineering and SRE teams; 5+ years cloud engineering experience with 2+ years managing teams; experience with Terraform/CloudFormation, Grafana/Datadog, GitHub Actions/CircleCI; Agile and strong communication.
Radar: Infrastructure for geofencing, mapping, and location-based fraud detection.
Experience managing SRE or production infrastructure teams, Terraform and AWS/EKS proficiency, high-availability and multi-region architecture experience, MongoDB sharded clusters, incident/observability expertise, customer-facing communication skills.
Cockroach Labs: Develops a distributed SQL database for cloud-native applications
Experience leading global operations and incident management; strong SRE/Production Engineering background; distributed systems and cloud infra knowledge; Go and Python proficiency; AI in operations; leadership and coaching of engineers; cross-team collaboration.