121 systems reliability engineer jobs at 106 companies in Cresskill, NJ

1w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
3w
Save
Mark Applied
Hide
BAE Systems - Principal Reliability Engineer
Wayne, New Jersey, United States
OnsiteFull Time
BAE Systems
BAE SystemsLondon Stock Exchange: BA: Provides advanced defense, aerospace, and security technology solutions.
6+ YOEBachelor's in engineering and 6+ years experience (4 with MS); strong reliability, failure analysis, statistical skills; proficiency with Microsoft Office; ability to obtain security clearance.
Microsoft Excel, Microsoft Word, Microsoft PowerPoint, Minitab, JMP, Python, FRACAS
1mo
Save
Mark Applied
Hide
Reliability Engineer
New York City, New York, United States
$165k-$250k/yr HybridFull Time
Two Sigma
Two Sigma: Systematic investment management and quantitative trading firm.
1+ YOE1+ years reliability engineering experience (5+ preferred), BS/BA in Computer Science or technical discipline, proficiency in Python/Java/C++/Rust, experience with automation, UNIX/Linux and distributed systems preferred.
Python, Java, C++, Rust, UNIX, Linux
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Jersey City or McLean or Richmond
HybridFull Time
Exiger
Exiger: AI-powered supply chain risk and compliance management software.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
AWS, Codex, Claude, Chaos Monkey, Gremlin, LitmusChaos, Snowflake, Redshift, Apache Iceberg, Go, C, Java, Linux/Unix
1mo
Save
Mark Applied
Hide
Staff Quality & Reliability Engineer
San Francisco or New York City
HybridFull Time
Beast Industries
Beast Industries: Produces digital media and consumer goods for MrBeast brands.
Expert in software quality engineering and site reliability for consumer-scale distributed systems; owns test strategy, SLOs/error budgets, CI/CD test gates, observability, incident response, and reliability tooling.
CI/CD
1w
Save
Mark Applied
Hide
Electrical Reliability Studies Engineer
New York or Los Angeles or Chicago or Houston or Tempe or Philadelphia or Dallas or North Miami Beach or Denver
$111k-$145k/yr HybridFull Time
Jacobs
JacobsNYSE: J: Global provider of professional engineering and technical services.
4+ YOEBachelor's in electrical engineering, 4+ years power-system and reliability analysis experience, familiarity with reliability indicators and NERC/FERC standards, strong analytical and communication skills.
ETAP, ReliaSoft BlockSim, Isograph Availability Workbench, Python
3w
Save
Mark Applied
Hide
Site Reliability Engineer
New York, New York, United States
HybridFull Time
Mistral AI
Mistral AI: Developing frontier artificial intelligence models and enterprise AI solutions.
7+ YOE7+ years SRE/DevOps experience, Master’s in CS/Engineering or related, strong cloud and distributed systems skills, Kubernetes/CI-CD/infra-as-code proficiency, scripting experience, observability and on-call experience.
Kubernetes, Flux, Terraform, Docker, Prometheus, Grafana, ELK Stack, Datadog, CloudFormation, Python, Go, Bash, Slurm, Fluidstack, Coreweave, Vast
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or United States or Canada
RemoteFull Time
Sanity
Sanity: Cloud platform for building and managing structured content-driven applications.
5+ YOE5+ years SRE on-call experience; strong experience with Kubernetes, GCP, observability (Prometheus), CI/CD, high-volume distributed systems, troubleshooting, and mentoring engineers.
Kubernetes, Prometheus, ElasticSearch, PostgreSQL, NATS, Kong, Fastly, Google Cloud Platform
2w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
New York City or Toronto
$184k-$240k/yr OnsiteFull Time
Movable Ink
Movable Ink: Provides AI-powered content personalization for digital marketing campaigns.
6+ YOE6+ years SRE/Software Engineering experience designing and operating scalable, multi-cloud distributed systems; expertise with observability, IaC, Kubernetes, and multiple programming languages.
Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB, Cassandra, Prometheus, Thanos, Grafana Alloy, Tempo, Terraform, Chef, EKS, GKE, NodeJS, Golang, Ruby, Python, shell
1mo
Save
Mark Applied
Hide
Senior Platform Reliability Engineer
San Francisco or New York City or Seattle
$182k-$250k/yr HybridFull Time
Grow Therapy
Grow Therapy: Platform connecting mental health providers with patients and insurance.
6+ YOE6+ years operating production systems; hands-on AWS, Kubernetes (EKS), Terraform; experience defining SLOs/SLAs and observability (DataDog); strong communication and systems-thinking skills; PostgreSQL experience a plus.
AWS, Kubernetes, EKS, Terraform, DataDog, PostgreSQL, Gem
1d
Save
Mark Applied
Hide
Senior Lead Platform Reliability Engineer
Charlotte or Irving or Chandler or West Des Moines or Iselin
$159k-$305k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
7+ YOE7+ years systems engineering or architecture, 5+ years supporting enterprise production environments, deep expertise in one infrastructure domain, SRE practices, automation and troubleshooting across domains.
Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, Python, Bash, PowerShell, Git, Ansible, Terraform, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Tech Infrastructure - USDS
New York, New York, United States
$137k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
3+ YOE3+ years SRE/systems engineering experience, bachelor\u0002s in CS or related, proficiency in Python/Go/Java/Shell, Linux, cloud and distributed systems, monitoring and incident management.
Python, Go, Java, Shell, Linux, Docker, Kubernetes, Prometheus, Grafana
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Chicago or New York City
$159k-$191k/yr HybridFull Time
Wonder
Wonder: Provides a multi-restaurant food delivery and meal kit platform.
Hands-on IaC (Terraform/Pulumi), CI/CD (Jenkins/Spinnaker), Docker, Linux systems, AWS and GCP, observability (Vector/Datadog/Splunk), Python/Bash scripting, security for payment environments, and mentoring experience.
Vector, Datadog, Splunk, AWS, GCP, Jenkins, Spinnaker, Docker, Terraform, Pulumi, AMI, Linux, Python, Bash, Ansible, EKS, GKE, Kubernetes
1mo
Save
Mark Applied
Hide
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Apache Airflow, AWS, Azure, CI/CD, GCP, Infrastructure as Code (IaC), Kubernetes, Linux, Python
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Chicago or New York City or Denver
$159k-$172k/yr HybridFull Time
Grubhub
Grubhub: A technology that connects diners with local restaurants through an online ordering and delivery platform.
Expertise in infrastructure automation, IaC (Terraform/Pulumi), CI/CD (Jenkins/Spinnaker), containerization (Docker), Linux systems, Python/Bash scripting, AWS and GCP, observability (Vector, Datadog, Splunk), and security for payment environments.
Terraform, Pulumi, Jenkins, Spinnaker, Docker, Amazon EC2, AMI, Linux, Python, Bash, AWS, GCP, Vector, Datadog, Splunk, Ansible, Kubernetes, EKS, GKE
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Berkeley Heights or Alpharetta or Sunnyvale
$128k-$216k/yr OnsiteFull Time
Fiserv
FiservNew York Stock Exchange: FI: Provides financial technology and payment processing services to institutions.
5+ YOE5+ years production experience with AWS, Kubernetes, and Linux; strong Terraform, CI/CD (GitHub Actions), Docker, GitHub, RDBMS/Document storage, and scripting (Python/Bash/Node/Ruby); experience designing scalable cloud systems.
Amazon Web Services, Kubernetes, GitHub Actions, Terraform, New Relic, Dynatrace, Datadog, Docker, GitHub, Python, Bash, Node, Ruby on Rails
1mo
Save
Mark Applied
Hide
Senior Systems Engineer
New York, New York, United States
OnsiteFull Time
Carrier
CarrierNYSE: CARR: Manufactures HVAC, refrigeration, and fire safety systems.
3+ YOEBachelor's degree, 3+ years experience in refrigeration/HVAC product design and development, expertise in heat transfer, fan systems, thermal management, testing, and reliability analysis.
SAP, Creo, Ansys/Fluent, Matlab / Simulink, Visual Basic, C/C++
2d
Save
Mark Applied
Hide
Sr Software Engineer - Reliability Engineering
North Hills, New York, United States
$122k-$203k/yr HybridFull Time
Cox Enterprises
Cox Enterprises: Providing global communications, automotive services, and media solutions.
5+ YOE5+ years in software/platform/infrastructure engineering, strong coding (Python/Go/Java), AWS and Terraform experience, SRE and observability knowledge, system design and incident response skills.
Python, Go, Java, AWS EC2, AWS RDS, AWS DynamoDB, AWS S3, AWS Aurora, AWS Lambda, AWS VPCs, AWS Athena, Terraform, Docker, Kubernetes, Linux, Windows, New Relic, Splunk, Prometheus, CI/CD pipelines
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE)
New York City, New York, United States
$220k-$350k/yr OnsiteFull Time
DualEntry
DualEntry: AI-native ERP software for automated enterprise financial operations.
5+ YOE5+ years experience with scalable systems, 3+ years Terraform, AWS, Python, observability (OpenTelemetry), CI/CD, on-call familiarity, strong ownership and problem-solving.
Terraform, OpenTelemetry, AWS, Python, GitHub, Postgres