171 site reliability engineer jobs at 98 companies in Secaucus, NJ

1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Jersey City or McLean or Richmond
HybridFull Time
Exiger
Exiger: Private supply-chain AI software serving corporations, government agencies, and banks with risk and compliance technology.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
AWS, Codex, Claude, Chaos Monkey, Gremlin, LitmusChaos, Snowflake, Redshift, Apache Iceberg, Go, C, Java, Linux/Unix
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
Braze
BrazeNASDAQ: BRZE: Customer engagement platform for cross-channel marketing and analytics.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Ruby on Rails, Ruby, Go, Linux, Unix Shell, Docker, Kubernetes, Terraform, Chef, MongoDB, Redis, Kafka, Postgres, PagerDuty
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
New York, New York, United States
HybridFull Time
Mistral AI
Mistral AI: Developer of open-weight and frontier AI models.
7+ YOE7+ years SRE/DevOps experience, Master’s in CS/Engineering or related, strong cloud and distributed systems skills, Kubernetes/CI-CD/infra-as-code proficiency, scripting experience, observability and on-call experience.
Kubernetes, Flux, Terraform, Docker, Prometheus, Grafana, ELK Stack, Datadog, CloudFormation, Python, Go, Bash, Slurm, Fluidstack, Coreweave, Vast
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Global financial services and banking institution.
4+ YOE4+ years cloud/platform engineering experience with GCP exposure, Terraform/IaC, CI/CD, observability, DevSecOps practices, incident response, and automation for reliability and resiliency.
GCP, Azure, Terraform, Terraform Enterprise, Log Analytics, Dynatrace, Resource Graph, CI/CD, IAM
1d
Save
Mark Applied
Hide
Site Reliability Engineer
Ontario or Canada or New York City or Ireland
$88k-$127k/yr RemoteFull Time
Greenhouse
Greenhouse: Private hiring software helping organizations source, interview, and onboard candidates with structured, AI-powered recruiting tools.
3+ YOE3+ years in site reliability or infrastructure, production AWS and Kubernetes experience, software delivery experience, and strong fluency with AI coding tools such as Claude. Must be eligible to work in Canada.
AWS, Kubernetes, RDS, Redis, OpenSearch, Claude
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Hyderabad or New York City or Chicago or London or Singapore or Tokyo or Hong Kong or Europe or United States or Asia-Pacific
HybridFull Time
Pico
Pico: Private financial-markets technology providing trading infrastructure, connectivity, market data, software, and analytics to institutions.
Bachelor's degree or relevant experience; financial markets technology experience; Linux, networking, computer architecture, programming or scripting, customer service, communication, and collaborative teamwork skills.
Linux, Python, C, C++, Java
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Greenwich, Connecticut, United States
HybridFull Time
Interactive Brokers
Interactive BrokersNASDAQ: IBKR: Global electronic brokerage firm providing automated trading technology.
5+ YOE5+ years experience in Linux/Unix, networking and coding; experience with cloud (AWS or Azure), Terraform or CloudFormation, Docker and Kubernetes; bachelor's or master's in CS/STEM; CI/CD, on-call rotation, mentoring skills.
CI/CD, Terraform, CloudFormation, AWS, Azure, Docker, Kubernetes, Linux, Unix
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Global financial services and banking institution.
4+ YOE4+ years cloud/platform engineering experience with Terraform and GCP; strong IaC, observability, automation, incident response, and DevSecOps skills; ability to define SLIs/SLOs and mentor engineers.
Terraform, Terraform Enterprise, Google Cloud Platform (GCP), Azure, Log Analytics, Dynatrace, Resource Graph, CI/CD
2w
Save
Mark Applied
Hide
Platform / Site Reliability Engineer
New York City, New York, United States
OnsiteFull Time
Sunset
Sunset: Private startup wind-down service helping founders close companies through legal, tax, and operational work.
Production cloud infrastructure and reliability experience across multiple services, strong software engineering skills, infrastructure and application coding, incident leadership, recovery expertise, and AI engineering tool proficiency.
AWS, Terraform, CI/CD, SOC 2
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
New York City or Toronto
$184k-$240k/yr OnsiteFull Time
Movable Ink
Movable Ink: AI-powered marketing software helping marketers personalize customer experiences across email, mobile, and web.
6+ YOE6+ years SRE/Software Engineering experience designing and operating scalable, multi-cloud distributed systems; expertise with observability, IaC, Kubernetes, and multiple programming languages.
Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB, Cassandra, Prometheus, Thanos, Grafana Alloy, Tempo, Terraform, Chef, EKS, GKE, NodeJS, Golang, Ruby, Python, shell
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or United States or Canada
RemoteFull Time
Sanity
Sanity: The Content Operating System for the AI era.
5+ YOE5+ years SRE on-call experience; strong experience with Kubernetes, GCP, observability (Prometheus), CI/CD, high-volume distributed systems, troubleshooting, and mentoring engineers.
Kubernetes, Prometheus, ElasticSearch, PostgreSQL, NATS, Kong, Fastly, Google Cloud Platform
4w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Jersey City, New Jersey, United States
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services and investment banking firm.
5+ YOE5+ years applied SRE experience, proficiency in reliability, scalability, observability, CI/CD, containers, networking, and mentoring; fluent in at least one of Python, Java/Spring Boot, or .Net; experience with enterprise AI in SRE workflows.
Python, Java, Spring Boot, .Net, CI/CD
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yr HybridFull Time
Radix Health
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
AI, HIPAA, PHI, SOC 2, 401(k)
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Pragma
New York, New York, United States
$175k-$230k/yr HybridFull Time
MarketAxess
MarketAxessNASDAQ: MKTX: Public financial-technology operating an electronic fixed-income trading platform for institutional investors and broker-dealers.
Experience with Java/Python, shell scripting, Linux, SQL, CI/CD (Jenkins), FIX, containers, AWS, monitoring, SRE practices, automation, and strong communication; Bachelor\u0002s/Master\u0002s in CS/Engineering or related field.
Java, Python, bash, ksh, JVM, Linux, SQL, Jenkins, FIX, AWS, GitHub Copilot
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Austin or New York City
$152k-$195k/yr HybridFull Time
SecurityScorecard
SecurityScorecard: Private cybersecurity ratings and third-party risk platform serving organizations managing supply-chain risk.
6+ YOE6+ years in SRE/DevOps with production Kubernetes, CI/CD pipeline expertise, IaC (Terraform/Helm/Pulumi), Python/Bash/Go proficiency, observability tooling, and experience with Kafka/Flink/ClickHouse and AI/LLM tooling integration.
Kubernetes, MCP servers, CI/CD, GitHub Actions, Jenkins, GitLab CI, EKS, GKE, AKS, Terraform, Helm, Argo CD, Pulumi, GitOps, Python, Bash, Go, Prometheus, Grafana, Datadog, OpenTelemetry, Kafka, Flink, ClickHouse, Langsmith, Langfuse
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute
San Francisco or New York or Austin or Seattle
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experience owning large GPU/compute fleets, automation of repair/deployment pipelines, firmware/BMC/Redfish familiarity, incident response and paging, observability and metrics tooling, and proficiency with production automation.
Redfish, BMC, IPMI, Temporal, Cadence, Prometheus, Grafana, Go, Python, Kubernetes, Claude Code, Cursor, LLM APIs, MCP servers
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York, New York, United States
$178k-$258k/yr RemoteFull Time
Adobe
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
Requires a computer science bachelor's degree or equivalent experience, Python, production ML inference, AWS cloud infrastructure, Kubernetes, vulnerability management, distributed-systems debugging, and on-call participation.
Python, PHP, Node.js, Ruby, SageMaker, OpenAI, Bedrock, EC2 Auto Scaling Groups, Kubernetes, AWS, Azure, GCP, AMI, LangGraph, LLM gateway, MCP, Aurora PostgreSQL, Memcached, Terraform, Terragrunt, Atlantis, Chef, Ansible, SSM, Docker, bash, Jenkins, Argo CD, New Relic, Splunk, Grafana, Prometheus, Fastly, Datadome, WAF, vLLM, LangSmith, Lambda
1mo
Save
Mark Applied
Hide
Staff Engineer, Site Reliability
New York City or Los Angeles or United States or Canada
RemoteFull Time
Babylist
Babylist: Private baby-registry and e-commerce platform helping expecting parents plan, shop, and prepare.
Hands-on Terraform expertise, proven AWS experience (EKS, RDS, networking, CDNs), production Kubernetes operation, CI/CD design, observability and alerting, on-call/incident management, cross-functional developer support, familiarity with AI tooling.
Ruby on Rails, AWS, Sidekiq, MySQL, Redis, Terraform, EKS, RDS, Kubernetes, CircleCI, GitHub Actions, Datadog, Sentry, PagerDuty, Cronitor, Claude, ChatGPT
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Chicago or New York City or Denver
$159k-$172k/yr HybridFull Time
Grubhub
Grubhub: U.S. food-ordering and delivery marketplace connecting diners with local restaurants, merchants, and convenience retailers.
Expertise in infrastructure automation, IaC (Terraform/Pulumi), CI/CD (Jenkins/Spinnaker), containerization (Docker), Linux systems, Python/Bash scripting, AWS and GCP, observability (Vector, Datadog, Splunk), and security for payment environments.
Terraform, Pulumi, Jenkins, Spinnaker, Docker, Amazon EC2, AMI, Linux, Python, Bash, AWS, GCP, Vector, Datadog, Splunk, Ansible, Kubernetes, EKS, GKE
1mo
Save
Mark Applied
Hide
Site Reliability / Infrastructure Engineer
New York City, New York, United States
$180k-$275k/yr OnsiteFull Time
Medal
Medal: Private AI research lab building action and world models for virtual and physical environments.
Experienced with Terraform, Elasticsearch, GCP/Kubernetes, relational database scaling (MySQL/Postgres), incident response, GitHub Actions; strong communication and startup experience preferred.
Terraform, Elasticsearch, GCP, Kubernetes, VPC, IAM, Cloud Logging, MySQL, Postgres, GitHub Actions, CircleCI, Salt, Redis, RabbitMQ, Electron, React, Redux, Styled Components, C#, C++, Swift, Kotlin, Java