109 platform reliability engineer jobs at 75 companies in Berkeley Heights, NJ

1w
Save
Mark Applied
Hide
Platform / Site Reliability Engineer
New York City, New York, United States
OnsiteFull Time
Sunset
Sunset: Handles legal and operational tasks for winding down startups.
Production cloud infrastructure and reliability experience across multiple services, strong software engineering skills, infrastructure and application coding, incident leadership, recovery expertise, and AI engineering tool proficiency.
AWS, Terraform, CI/CD, SOC 2
2mo
Save
Mark Applied
Hide
Senior Platform Reliability Engineer
San Francisco or New York City or Seattle
$182k-$250k/yr HybridFull Time
Grow Therapy
Grow Therapy: Platform connecting mental health providers with patients and insurance.
6+ YOE6+ years operating production systems; hands-on AWS, Kubernetes (EKS), Terraform; experience defining SLOs/SLAs and observability (DataDog); strong communication and systems-thinking skills; PostgreSQL experience a plus.
AWS, Kubernetes, EKS, Terraform, DataDog, PostgreSQL, Gem
6h
Save
Mark Applied
Hide
Platform Reliability Engineer
Ramsey, New Jersey, United States
$127k-$192k/yr OnsiteFull Time
Konica Minolta
Konica MinoltaTokyo Stock Exchange: 4902: Provides digital workplace solutions, imaging technology, and IT services.
8+ YOERequires 8+ years administering enterprise databases and middleware, including Microsoft SQL Server, Linux/RHEL, WebSphere, OpenLDAP, SSL/TLS, DNS, scripting, and incident resolution. Bachelor's preferred.
Microsoft SQL Server, SSIS, SQL Agent, T-SQL, Linux, RHEL, Apache, PHP, SSL/TLS, Java keystores, DNS, IBM WebSphere Application Server (WAS), OpenLDAP, Windows Server, RDP, PowerShell, bash, IBM DB2, SSRS, Azure Data Factory, Power BI, Tidal, Microsoft Entra ID, Uptrends, MuleSoft, SAP, Salesforce, Ansible, Terraform, PowerShell DSC, Microsoft Azure, Microsoft 365, GitHub Copilot, Datadog, Dynatrace, Docker, OpenLiberty, Ollama, vLLM, REST, API, Tenable, Bitsight
3w
Save
Mark Applied
Hide
Staff , Site Reliability Engineer - Cloud Platform
New York City, New York, United States
$190k-$210k/yr HybridFull Time
Butterfly Network
Butterfly NetworkNYSE: BFLY: Handheld whole-body ultrasound scanners powered by semiconductor technology.
8+ YOE8+ years managing production systems; deep hands-on AWS and Kubernetes experience; strong programming/scripting and automation skills; observability platform ownership; incident response leadership and mentoring ability.
Kubernetes, EKS, AWS, NewRelic, Datadog, DICOM, HL7, FHIR, PACS, VNA, EMR
2mo
Save
Mark Applied
Hide
Director of Platform & Reliability Engineering
San Francisco or New York City
$235k-$245k/yr HybridFull Time
Forge Global
Forge GlobalNYSE: FRGE: Marketplace for trading private shares and pre-IPO stock.
8+ YOE5+ Mgmt8+ years software engineering experience with infrastructure/platform focus, 5+ years people leadership, deep cloud/observability/incident response experience, strong distributed systems judgment, and ability to set platform strategy.
Kubernetes, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Jersey City or McLean or Richmond
HybridFull Time
Exiger
Exiger: AI-powered supply chain risk and compliance management software.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
AWS, Codex, Claude, Chaos Monkey, Gremlin, LitmusChaos, Snowflake, Redshift, Apache Iceberg, Go, C, Java, Linux/Unix
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Provides global banking, investing, and financial risk management services.
4+ YOE4+ years cloud/platform engineering experience with GCP exposure, Terraform/IaC, CI/CD, observability, DevSecOps practices, incident response, and automation for reliability and resiliency.
GCP, Azure, Terraform, Terraform Enterprise, Log Analytics, Dynatrace, Resource Graph, CI/CD, IAM
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
New York City or United States
$135k-$160k/yr RemoteFull Time
Fabric
Fabric: Provides clinical automation and care enablement software for healthcare.
5+ YOE5+ years SRE or platform engineering experience with AWS/EKS, production Kubernetes, Terraform, Datadog, Helm, GitHub Actions, and coding in Python/Bash/Go; HIPAA compliance experience preferred.
AWS, EKS, EC2, RDS, S3, Kubernetes (EKS), Terraform, Datadog, Helm, GitHub Actions, Python, Bash, Go
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Provides banking, investment, and financial risk management services.
4+ YOE4+ years cloud/platform engineering experience with Terraform and GCP; strong IaC, observability, automation, incident response, and DevSecOps skills; ability to define SLIs/SLOs and mentor engineers.
Terraform, Terraform Enterprise, Google Cloud Platform (GCP), Azure, Log Analytics, Dynatrace, Resource Graph, CI/CD
3mo
Save
Mark Applied
Hide
Cloud Site Reliability Engineer (SRE) - Data Management & Analytics Platform
Princeton, New Jersey, United States
HybridFull Time
Bloomberg
Bloomberg: Delivers financial data, news, and software to global markets.
5+ YOE5+ years in SRE/DevOps/Cloud Infra; Python and/or Go; AWS; IaC; data platforms experience is a plus.
Prometheus, Grafana, CloudWatch, Datadog, Terraform, CloudFormation, Docker, Kubernetes, AWS S3, EMR, Kinesis, Glue, Redshift, Databricks, Snowflake
2w
Save
Mark Applied
Hide
Staff Platform Database Reliability Engineer
Toronto or New York City or Montreal or Kitchener or London or San Francisco or North America
HybridFull Time
Index Exchange
Index Exchange: Provides a programmatic marketplace for digital advertising transactions.
7+ YOERequires 7+ years in database reliability, database administration, SRE, or similar roles; deep database expertise; Kubernetes and bare-metal operations; automation, observability, troubleshooting, leadership, and communication skills.
MySQL, MariaDB, Galera, PostgreSQL, Aerospike, Redis, Kafka, Zookeeper, Cassandra, ScyllaDB, etcd, StarRocks, Vertica, Trino, Iceberg, Kubernetes, Ansible, Terraform, Grafana, Prometheus, Loki, Python, Bash, Shell, Golang, Kotlin, Java, Helm, GitLab, ArgoCD, Liquibase, Flyway, Alembic, Spark, Flink, Hadoop, Presto
1d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yr HybridFull Time
Pivotal Health
Pivotal Health: AI platform automating healthcare insurance claim disputes for providers.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
AI, HIPAA, PHI, SOC 2, 401(k)
4w
Save
Mark Applied
Hide
Member of Technical Staff, Platform Engineer
New York City, New York, United States
$200k-$300k/yr OnsiteFull Time
Arca
Arca: AI-native wealth management platform for personalized financial advice.
Experienced platform engineer to own infrastructure, backend systems, developer experience, sandboxing for agents, reliability, performance, and M&A onboarding at scale.
2mo
Save
Mark Applied
Hide
Senior VoIP Operations & Reliability Engineer (Carrier-Class Voice Platform)
Newton, New Jersey, United States
OnsiteFull Time
Planet Networks
Planet Networks: Provides fiber optic internet and telecommunications infrastructure services.
Senior hands-on experience operating carrier-scale VoIP systems (SIP, Kamailio/OpenSIPS, Asterisk), reliability engineering, incident response, SLOs/SLIs, observability, and Linux automation.
Kamailio, Asterisk, OpenSIPS, Homer, HEP, sipp, sngrep, Wireshark, Prometheus, Grafana, Python, Lua, shell, CI/CD, PJSIP, ARI, AMI, FreeSWITCH, RADIUS, Diameter
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Observability
Chicago or New York City
$160k-$200k/yr HybridFull Time
Ripple
Ripple: Provides blockchain solutions for global payments and liquidity.
7+ YOE7+ years SRE/Platform experience focused on observability, New Relic, Terraform, PowerShell, Azure/AWS, incident management (Incident.IO/PagerDuty/OpsGenie), and coaching engineering teams.
New Relic, NRQL, Terraform, PowerShell, Azure, AWS, Azure DevOps, Octopus Deploy, Incident.IO, PagerDuty, OpsGenie, Slack, Python, Bash, Jira, SQL Server
4w
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
1mo
Save
Mark Applied
Hide
The Better Money Company - Platform Engineer
New York, New York, United States
OnsiteFull Time
YieldNest
YieldNest: Liquid restaking protocol for risk-adjusted DeFi yields.
3+ YOE3+ years in platform/infra or reliability engineering; experience building test infrastructure for payments/ledgers; familiarity with formal verification, property-based or chaos testing preferred; strong ownership of internal tooling and CI/CD.
3w
Save
Mark Applied
Hide
Sr Software Engineer - Reliability Engineering
North Hills, New York, United States
$122k-$203k/yr HybridFull Time
Cox Enterprises
Cox Enterprises: Providing global communications, automotive services, and media solutions.
5+ YOE5+ years in software/platform/infrastructure engineering, strong coding (Python/Go/Java), AWS and Terraform experience, SRE and observability knowledge, system design and incident response skills.
Python, Go, Java, AWS EC2, AWS RDS, AWS DynamoDB, AWS S3, AWS Aurora, AWS Lambda, AWS VPCs, AWS Athena, Terraform, Docker, Kubernetes, Linux, Windows, New Relic, Splunk, Prometheus, CI/CD pipelines
2w
Save
Mark Applied
Hide
Service Reliability Engineer
London or Manchester or New York City
HybridFull Time
Fitch Group
Fitch Group: Provides global credit ratings and financial market research services.
Deep SRE, DevOps, or platform engineering experience with AWS, Azure, Docker, Kubernetes, Linux, Windows, CI/CD, cloud security, networking, and Python, PowerShell, or Bash.
AWS, Azure, Docker, Kubernetes, Linux, Windows, IIS, .NET, Java Spring Boot, GitHub Actions, Bamboo, Python, PowerShell, Bash, Datadog, Microsoft Teams, AWS Bedrock, SageMaker, Model Context Protocol (MCP), IAM, OPA, AWS Config, AWS CloudTrail, AWS Security Hub, Wiz, DNS, CIS, NIST, ISO 27001
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Release Engineering
New York, New York, United States
$208k-$274k/yr HybridFull Time
Plaid
Plaid: Provides financial data connectivity and payment infrastructure via APIs.
8+ YOE8+ years in backend/SRE/platform engineering; experience designing SLO/SLI programs, progressive delivery, canary rollouts, metric-gated analysis, and automated rollback; proficiency in Go or similar; familiarity with Kubernetes, Prometheus, ArgoCD; strong leadership and incident response skills.
Go, Kubernetes, Prometheus, ArgoCD