132 site reliability engineer jobs at 89 companies in Princeton, NJ

1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Jersey City or McLean or Richmond
HybridFull Time
Exiger
Exiger: AI-powered supply chain risk and compliance management software.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
AWS, Codex, Claude, Chaos Monkey, Gremlin, LitmusChaos, Snowflake, Redshift, Apache Iceberg, Go, C, Java, Linux/Unix
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
Braze
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Ruby on Rails, Ruby, Go, Linux, Unix Shell, Docker, Kubernetes, Terraform, Chef, MongoDB, Redis, Kafka, Postgres, PagerDuty
4w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yr RemoteFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
New York City or United States
$135k-$160k/yr RemoteFull Time
Fabric
Fabric: Provides clinical automation and care enablement software for healthcare.
5+ YOE5+ years SRE or platform engineering experience with AWS/EKS, production Kubernetes, Terraform, Datadog, Helm, GitHub Actions, and coding in Python/Bash/Go; HIPAA compliance experience preferred.
AWS, EKS, EC2, RDS, S3, Kubernetes (EKS), Terraform, Datadog, Helm, GitHub Actions, Python, Bash, Go
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
New York, New York, United States
HybridFull Time
Mistral AI
Mistral AI: Developing frontier artificial intelligence models and enterprise AI solutions.
7+ YOE7+ years SRE/DevOps experience, Master’s in CS/Engineering or related, strong cloud and distributed systems skills, Kubernetes/CI-CD/infra-as-code proficiency, scripting experience, observability and on-call experience.
Kubernetes, Flux, Terraform, Docker, Prometheus, Grafana, ELK Stack, Datadog, CloudFormation, Python, Go, Bash, Slurm, Fluidstack, Coreweave, Vast
1w
Save
Mark Applied
Hide
Site Reliability Engineer
North Carolina or King of Prussia or United States
$130k-$160k/yr RemoteFull Time
Qlik
Qlik: Provides data integration and analytics software for business intelligence.
3+ YOEBachelor's in CS or related,3+ years cloud engineering with AWS/Azure,Kubernetes,Infrastructure as Code (Terraform/Crossplane/Ansible),microservices,scripting (Bash/Python/Go),observability,incident/on-call experience,and ability to obtain IL5 clearance;US citizenship requirement.
Terraform, Crossplane, Ansible, Kubernetes, AWS, Azure, Bash, Python, Go, CI/CD, Vault, AWS SSM, Prometheus, Open Telemetry, SIEM, Splunk, Helm, Temporal, Clik House, Fire Hydrant, Grafana, Solace, Gloo
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Provides global banking, investing, and financial risk management services.
4+ YOE4+ years cloud/platform engineering experience with GCP exposure, Terraform/IaC, CI/CD, observability, DevSecOps practices, incident response, and automation for reliability and resiliency.
GCP, Azure, Terraform, Terraform Enterprise, Log Analytics, Dynatrace, Resource Graph, CI/CD, IAM
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Camden, New Jersey, United States
OnsiteFull Time
Sporttrade
Sporttrade: A sports betting exchange for trading sports outcomes.
5+ YOE5+ years SRE/DevOps experience supporting 24/7 production systems; strong Linux and scripting; Kubernetes, Terraform, Ansible, Jenkins; observability (Datadog/Prometheus/Grafana); PostgreSQL and Kafka experience; Java production debugging and networking skills.
Datadog, Wazuh, Prometheus, Grafana, Kubernetes, GKE, EKS, Istio, AWS, GCP, Ansible, Terraform, Jenkins, HashiCorp Vault, PostgreSQL, Kafka, Confluent Cloud, Redis, Java
1w
Save
Mark Applied
Hide
Site Reliability Engineer
Hyderabad or New York City or Chicago or London or Singapore or Tokyo or Hong Kong or Europe or United States or Asia-Pacific
HybridFull Time
Pico
Pico: Provides managed infrastructure and data services to financial markets.
Bachelor's degree or relevant experience; financial markets technology experience; Linux, networking, computer architecture, programming or scripting, customer service, communication, and collaborative teamwork skills.
Linux, Python, C, C++, Java
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Provides banking, investment, and financial risk management services.
4+ YOE4+ years cloud/platform engineering experience with Terraform and GCP; strong IaC, observability, automation, incident response, and DevSecOps skills; ability to define SLIs/SLOs and mentor engineers.
Terraform, Terraform Enterprise, Google Cloud Platform (GCP), Azure, Log Analytics, Dynatrace, Resource Graph, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or United States or Canada
RemoteFull Time
Sanity
Sanity: Cloud platform for building and managing structured content-driven applications.
5+ YOE5+ years SRE on-call experience; strong experience with Kubernetes, GCP, observability (Prometheus), CI/CD, high-volume distributed systems, troubleshooting, and mentoring engineers.
Kubernetes, Prometheus, ElasticSearch, PostgreSQL, NATS, Kong, Fastly, Google Cloud Platform
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
New York City or Toronto
$184k-$240k/yr OnsiteFull Time
Movable Ink
Movable Ink: Provides AI-powered content personalization for digital marketing campaigns.
6+ YOE6+ years SRE/Software Engineering experience designing and operating scalable, multi-cloud distributed systems; expertise with observability, IaC, Kubernetes, and multiple programming languages.
Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB, Cassandra, Prometheus, Thanos, Grafana Alloy, Tempo, Terraform, Chef, EKS, GKE, NodeJS, Golang, Ruby, Python, shell
2w
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Jersey City, New Jersey, United States
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years applied SRE experience, proficiency in reliability, scalability, observability, CI/CD, containers, networking, and mentoring; fluent in at least one of Python, Java/Spring Boot, or .Net; experience with enterprise AI in SRE workflows.
Python, Java, Spring Boot, .Net, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Pragma
New York, New York, United States
$175k-$230k/yr HybridFull Time
MarketAxess
MarketAxessNASDAQ: MKTX: Operates an electronic trading platform for fixed-income securities.
Experience with Java/Python, shell scripting, Linux, SQL, CI/CD (Jenkins), FIX, containers, AWS, monitoring, SRE practices, automation, and strong communication; Bachelor\u0002s/Master\u0002s in CS/Engineering or related field.
Java, Python, bash, ksh, JVM, Linux, SQL, Jenkins, FIX, AWS, GitHub Copilot
1mo
Save
Mark Applied
Hide
Staff Engineer, Site Reliability
New York City or Los Angeles or United States or Canada
RemoteFull Time
Babylist
Babylist: Operates an e-commerce platform and universal registry for baby products.
Hands-on Terraform expertise, proven AWS experience (EKS, RDS, networking, CDNs), production Kubernetes operation, CI/CD design, observability and alerting, on-call/incident management, cross-functional developer support, familiarity with AI tooling.
Ruby on Rails, AWS, Sidekiq, MySQL, Redis, Terraform, EKS, RDS, Kubernetes, CircleCI, GitHub Actions, Datadog, Sentry, PagerDuty, Cronitor, Claude, ChatGPT
1mo
Save
Mark Applied
Hide
Site Reliability / Infrastructure Engineer
New York City, New York, United States
$180k-$275k/yr OnsiteFull Time
General Intuition
General Intuition: Developing AI models that perceive and act in dynamic environments.
Experienced with Terraform, Elasticsearch, GCP/Kubernetes, relational database scaling (MySQL/Postgres), incident response, GitHub Actions; strong communication and startup experience preferred.
Terraform, Elasticsearch, GCP, Kubernetes, VPC, IAM, Cloud Logging, MySQL, Postgres, GitHub Actions, CircleCI, Salt, Redis, RabbitMQ, Electron, React, Redux, Styled Components, C#, C++, Swift, Kotlin, Java
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Philadelphia, Pennsylvania, United States
HybridFull Time
FreedomPay
FreedomPay: A commerce technology that builds a global payment platform with observability, security, and automation capabilities.
5+ YOEBS in CS or equivalent; 5+ years in high-availability web environments; expert APM (Dynatrace/Datadog/New Relic); AI-driven ops experience; PowerShell/Python scripting, SQL/T-SQL, networking, container orchestration, Azure and VMware; on-call rotation experience.
Dynatrace, Datadog, New Relic, Anthropic (Claude), OpenAI (Codex), Azure AI services, Azure SRE Agent, Microsoft PowerShell, Python, SQL, T-SQL, Azure Kubernetes Service (AKS), Azure, VMware, Windows Server (IIS), PagerDuty Process Automation (Rundeck), DNS, HTTP/HTTPS, TCP/IP
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Chicago or New York City or Denver
$159k-$172k/yr HybridFull Time
Grubhub
Grubhub: A technology that connects diners with local restaurants through an online ordering and delivery platform.
Expertise in infrastructure automation, IaC (Terraform/Pulumi), CI/CD (Jenkins/Spinnaker), containerization (Docker), Linux systems, Python/Bash scripting, AWS and GCP, observability (Vector, Datadog, Splunk), and security for payment environments.
Terraform, Pulumi, Jenkins, Spinnaker, Docker, Amazon EC2, AMI, Linux, Python, Bash, AWS, GCP, Vector, Datadog, Splunk, Ansible, Kubernetes, EKS, GKE
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute
San Francisco or New York or Austin or Seattle
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Experience owning large GPU/compute fleets, automation of repair/deployment pipelines, firmware/BMC/Redfish familiarity, incident response and paging, observability and metrics tooling, and proficiency with production automation.
Redfish, BMC, IPMI, Temporal, Cadence, Prometheus, Grafana, Go, Python, Kubernetes, Claude Code, Cursor, LLM APIs, MCP servers
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Austin or New York City
$152k-$195k/yr HybridFull Time
SecurityScorecard
SecurityScorecard: Provides continuous cybersecurity ratings and risk monitoring for organizations.
6+ YOE6+ years in SRE/DevOps with production Kubernetes, CI/CD pipeline expertise, IaC (Terraform/Helm/Pulumi), Python/Bash/Go proficiency, observability tooling, and experience with Kafka/Flink/ClickHouse and AI/LLM tooling integration.
Kubernetes, MCP servers, CI/CD, GitHub Actions, Jenkins, GitLab CI, EKS, GKE, AKS, Terraform, Helm, Argo CD, Pulumi, GitOps, Python, Bash, Go, Prometheus, Grafana, Datadog, OpenTelemetry, Kafka, Flink, ClickHouse, Langsmith, Langfuse