83 platform reliability engineer jobs at 58 companies in Ridgefield, CT
1mo
Save
Mark Applied
Hide
1mo
Senior Platform Reliability Engineer
San Francisco or New York City or Seattle
$182k-$250k/yrHybridFull Time
Grow Therapy: Platform connecting mental health providers with patients and insurance.
6+ YOE6+ years operating production systems; hands-on AWS, Kubernetes (EKS), Terraform; experience defining SLOs/SLAs and observability (DataDog); strong communication and systems-thinking skills; PostgreSQL experience a plus.
8+ YOE8+ years managing production systems; deep hands-on AWS and Kubernetes experience; strong programming/scripting and automation skills; observability platform ownership; incident response leadership and mentoring ability.
Forge GlobalNYSE: FRGE: Marketplace for trading private shares and pre-IPO stock.
8+ YOE5+ Mgmt8+ years software engineering experience with infrastructure/platform focus, 5+ years people leadership, deep cloud/observability/incident response experience, strong distributed systems judgment, and ability to set platform strategy.
Exiger: AI-powered supply chain risk and compliance management software.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
Bank of AmericaNYSE: BAC: Provides global banking, investing, and financial risk management services.
4+ YOE4+ years cloud/platform engineering experience with GCP exposure, Terraform/IaC, CI/CD, observability, DevSecOps practices, incident response, and automation for reliability and resiliency.
Fabric: Provides clinical automation and care enablement software for healthcare.
5+ YOE5+ years SRE or platform engineering experience with AWS/EKS, production Kubernetes, Terraform, Datadog, Helm, GitHub Actions, and coding in Python/Bash/Go; HIPAA compliance experience preferred.
Bank of AmericaNYSE: BAC: Provides banking, investment, and financial risk management services.
4+ YOE4+ years cloud/platform engineering experience with Terraform and GCP; strong IaC, observability, automation, incident response, and DevSecOps skills; ability to define SLIs/SLOs and mentor engineers.
Arca: AI-native wealth management platform for personalized financial advice.
Experienced platform engineer to own infrastructure, backend systems, developer experience, sandboxing for agents, reliability, performance, and M&A onboarding at scale.
Ripple: Provides blockchain solutions for global payments and liquidity.
7+ YOE7+ years SRE/Platform experience focused on observability, New Relic, Terraform, PowerShell, Azure/AWS, incident management (Incident.IO/PagerDuty/OpsGenie), and coaching engineering teams.
New Relic, NRQL, Terraform, PowerShell, Azure, AWS, Azure DevOps, Octopus Deploy, Incident.IO, PagerDuty, OpsGenie, Slack, Python, Bash, Jira, SQL Server
YieldNest: Liquid restaking protocol for risk-adjusted DeFi yields.
3+ YOE3+ years in platform/infra or reliability engineering; experience building test infrastructure for payments/ledgers; familiarity with formal verification, property-based or chaos testing preferred; strong ownership of internal tooling and CI/CD.
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Santa Monica or Lower Manhattan or San Francisco or Los Angeles
$150k-$200k/yrHybridFull Time
Pivotal Health: AI platform automating healthcare insurance claim disputes for providers.
5+ YOE5+ years in platform, infrastructure, or software engineering; strong Python; cloud-native systems (GCP); Terraform; CI/CD; containers; event-driven architectures; security and reliability.
Cox Enterprises: Providing global communications, automotive services, and media solutions.
5+ YOE5+ years in software/platform/infrastructure engineering, strong coding (Python/Go/Java), AWS and Terraform experience, SRE and observability knowledge, system design and incident response skills.
Staff Site Reliability Engineer, Release Engineering
New York, New York, United States
$208k-$274k/yrHybridFull Time
Plaid: Provides financial data connectivity and payment infrastructure via APIs.
8+ YOE8+ years in backend/SRE/platform engineering; experience designing SLO/SLI programs, progressive delivery, canary rollouts, metric-gated analysis, and automated rollback; proficiency in Go or similar; familiarity with Kubernetes, Prometheus, ArgoCD; strong leadership and incident response skills.
Nscale: Vertically integrated AI infrastructure provider for high-performance computing.
Hands-on experience operating Kubernetes platforms, infrastructure automation and CI/CD/GitOps, strong Linux and networking fundamentals, production-quality automation in Go/Python/Bash, reliability engineering and mentoring skills.
Kubernetes, Go, Python, Bash, CI/CD, GitOps, Linux
MongoDBNASDAQ: MDB: Cloud-based document database platform for software application development.
6+ YOE6+ years software development and distributed systems experience; proficiency in Python or Go; experience building and operating large-scale CI/CD pipelines; Kubernetes and cloud platform expertise (AWS, GCP, Azure); Linux and networking knowledge; participate in 24/7 on-call.
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
10+ YOE5+ MgmtBachelor's in a technical field,10+ years engineering experience with 5+ years leading SRE/Platform teams; experience with observability, incident management, distributed systems, and cloud architecture.
AWS, New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, OpenTelemetry
Prague or Berlin or San Francisco or New York City
Kč70k-Kč85k/yrOnsiteFull Time
Zeta GlobalNYSE: ZETA: Provides an AI-powered marketing platform for enterprise customer engagement.
7+ YOE7+ years software engineering experience; strong Scala and Spark skills; experience with large-scale data-processing frameworks, Hadoop, Airflow, and AWS; mentoring and improving platform reliability.