76 infrastructure reliability engineer jobs at 67 companies in Cresskill, NJ

1mo
Save
Mark Applied
Hide
Customer Reliability Engineer - Infrastructure
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
Apache Airflow, AWS, Azure, CI/CD, GCP, Infrastructure as Code (IaC), Kubernetes, Linux, Python
2w
Save
Mark Applied
Hide
Reliability Engineer, R&D
Austin or New York City or San Francisco or Seattle
$203k-$232k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Tech Infrastructure - USDS
New York, New York, United States
$137k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
3+ YOE3+ years SRE/systems engineering experience, bachelor\u0002s in CS or related, proficiency in Python/Go/Java/Shell, Linux, cloud and distributed systems, monitoring and incident management.
Python, Go, Java, Shell, Linux, Docker, Kubernetes, Prometheus, Grafana
2mo
Save
Mark Applied
Hide
Lead Infrastructure Engineer
New York, New York, United States
$220k-$300k/yr RemoteFull Time
Bayview Asset Management
Bayview Asset Management: Investment management firm specializing in credit and mortgage assets.
5+ YOE2+ Mgmt5+ years building and operating production infrastructure; strong cloud and DevOps experience; hands-on leadership; expertise in IaC and platform reliability.
AWS, GCP, Azure, Terraform, Pulumi, CloudFormation, Bicep, Docker, Kubernetes, ECS, CI/CD, Git, Python, Go, TypeScript
5d
Save
Mark Applied
Hide
Platform Reliability Engineer - Principal Engineer
Iselin or Irving or Charlotte
$159k-$305k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Provides banking, investment, mortgage, and consumer finance products.
7+ YOE7+ years engineering experience, 5+ years supporting enterprise production environments, hands-on in one infrastructure domain, SRE practice experience, strong troubleshooting and automation skills.
Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, Python, Bash, PowerShell, Git, Ansible, Terraform, CICD
1mo
Save
Mark Applied
Hide
Staff Software Engineer, Infrastructure
Canada or United States or Seattle or Paris or New York City
$238k-$382k/yr RemoteFull Time
Docker
Docker: Provides a platform for building, sharing, and running containerized applications.
8+ YOE8+ years backend/infrastructure engineering; strong Go; experience with Kubernetes/EKS, cloud platforms, networking, and reliability; Bachelor's or equivalent; experience leading cross-team technical initiatives and strong written/verbal communication.
Go, Terraform, Argo CD, EKS, Envoy Gateway, Grafana Cloud, OpenTelemetry, Prometheus, Grafana, GitHub Actions, Kubernetes, Linux, GitOps, Covey Scout
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yr RemoteFull Time
Bedrock Ocean Exploration
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Python, Go, Bash, Terraform, AWS, Docker, Kubernetes, Prometheus, Grafana, ROS 2, ROS, Jetson, Linux, IAM
1d
Save
Mark Applied
Hide
Senior Lead Platform Reliability Engineer
Charlotte or Irving or Chandler or West Des Moines or Iselin
$159k-$305k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
7+ YOE7+ years systems engineering or architecture, 5+ years supporting enterprise production environments, deep expertise in one infrastructure domain, SRE practices, automation and troubleshooting across domains.
Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, Python, Bash, PowerShell, Git, Ansible, Terraform, CI/CD
1w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Software platform for building custom internal business applications.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Chicago or New York City or Denver
$159k-$172k/yr HybridFull Time
Grubhub
Grubhub: A technology that connects diners with local restaurants through an online ordering and delivery platform.
Expertise in infrastructure automation, IaC (Terraform/Pulumi), CI/CD (Jenkins/Spinnaker), containerization (Docker), Linux systems, Python/Bash scripting, AWS and GCP, observability (Vector, Datadog, Splunk), and security for payment environments.
Terraform, Pulumi, Jenkins, Spinnaker, Docker, Amazon EC2, AMI, Linux, Python, Bash, AWS, GCP, Vector, Datadog, Splunk, Ansible, Kubernetes, EKS, GKE
2w
Save
Mark Applied
Hide
Site Reliability Engineer II
Scottsdale or San Francisco or Chicago or New York City
$86k-$126k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
2+ YOEBachelor's or equivalent, minimum 2 years DevOps/Dev/SRE experience, Linux/Unix experience, infrastructure automation (Chef/Ansible/Puppet, Terraform), containerization (Docker,Kubernetes), cloud (AWS/GCP/Azure), on-call rotation.
Linux, Unix, Chef, Ansible, Puppet, Terraform, Docker, Kubernetes, AWS, GCP, Azure, Java, Ruby, Python, JavaScript, Go
2d
Save
Mark Applied
Hide
Sr Software Engineer - Reliability Engineering
North Hills, New York, United States
$122k-$203k/yr HybridFull Time
Cox Enterprises
Cox Enterprises: Providing global communications, automotive services, and media solutions.
5+ YOE5+ years in software/platform/infrastructure engineering, strong coding (Python/Go/Java), AWS and Terraform experience, SRE and observability knowledge, system design and incident response skills.
Python, Go, Java, AWS EC2, AWS RDS, AWS DynamoDB, AWS S3, AWS Aurora, AWS Lambda, AWS VPCs, AWS Athena, Terraform, Docker, Kubernetes, Linux, Windows, New Relic, Splunk, Prometheus, CI/CD pipelines
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Messaging Services
Secaucus or New York or New Jersey or United States
$140k-$150k/yr RemoteFull Time
NBA
NBA: Operates professional basketball leagues and manages global media rights.
10+ YOEBachelor's in computer science or related, 10+ years in enterprise messaging/infrastructure/reliability, deep Exchange Online/M365 and email security experience, PowerShell and Microsoft Graph automation, Slack/Teams support, incident response and executive support.
Microsoft Exchange Online (M365), Microsoft Outlook, SMTP, Proofpoint, PowerShell, Microsoft Graph, Slack, Microsoft Teams, DMARC, DKIM, SPF, ARC
2mo
Save
Mark Applied
Hide
Infrastructure Software Engineer
Zurich or United Kingdom or Germany or San Francisco or New York
HybridFull Time
Namespace
Namespace: Cloud infrastructure platform for faster software builds and tests.
Experience engineering large-scale infrastructure; strong network architecture and performance tuning; hands-on hardware deployment; proficient with orchestration and automation tools; track record of reliable systems and efficiency gains.
Go, Kubernetes, Terraform, Ansible
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer (SRE)
Chicago or New York City
$175k-$220k/yr HybridFull Time
Optimal Market Technologies
Optimal Market Technologies: Operates a broker-dealer platform for wholesale options execution.
Hands-on Linux and network administration, strong scripting/automation, Infrastructure-as-Code experience, production support and incident response ownership, staff management/mentoring, familiarity with C++, Python, SQL, Azure, and real-time trading systems.
C++, Python, SQL, Linux, CentOS7, RHEL9, Microsoft Azure, Microsoft Azure Virtual Desktop (AVD), Claude Code, FIX Protocol, PostgreSQL, AERON
2mo
Save
Mark Applied
Hide
SWE - Backend Infrastructure Engineer
San Francisco or Bellevue or New York
$175k-$280k/yr OnsiteFull Time
Sesame
Sesame: Designing wearable computers with lifelike voice-driven AI agents.
3+ YOEStrong systems thinker with reliability engineering experience; 3+ years in infrastructure, platform, or ML systems; Kubernetes production experience; strong communication.
Kubernetes, Terraform, CloudFormation, Pulumi, TorchServe, Seldon, KServe, Ray Serve, PyTorch, APIs, Database design
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (FedRAMP / Security) - NY
New York or Israel
$170k-$350k/yr RemoteFull Time
Coralogix
Coralogix: AI-powered observability and security data platform.
5+ YOE5+ years SRE/DevOps experience, strong Kubernetes and cloud (AWS) skills, infrastructure-as-code (Terraform/Crossplane), monitoring tooling (Prometheus/Grafana/Coralogix), Golang experience preferred, FedRAMP/compliance experience advantageous, must be located within EST/CT time zone.
Kubernetes, Kops, AWS, Kafka, Prometheus, Thanos, Coralogix, Git, Argo CD, Istio, Grafana, Terraform, Crossplane, Golang
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
New York City or United States
RemoteFull Time
Knock
Knock: Developer-first infrastructure platform for powering product notifications across channels.
4+ YOE4+ years DevOps/infra experience; Kubernetes with IaC; advanced AWS multi-account deployments; databases including Aurora Postgres, MongoDB, or ClickHouse; queues/streams (SQS, Kinesis, Kafka); strong reliability, scalability, and communication.
Terraform, Kubernetes, AWS, EKS, PostgreSQL, MongoDB, ClickHouse, Datadog, CloudWatch, Honeycomb, SQS, Kinesis, Kafka