58 reliability automation engineer jobs at 46 companies in Avenel, NJ

2mo
Save
Mark Applied
Hide
Reliability Engineer
New York City, New York, United States
$165k-$250k/yr HybridFull Time
Two Sigma Investment Management
Two Sigma Investment Management: Private quantitative investment manager using data science and technology to manage diversified global strategies for institutional investors.
1+ YOE1+ years reliability engineering experience (5+ preferred), BS/BA in Computer Science or technical discipline, proficiency in Python/Java/C++/Rust, experience with automation, UNIX/Linux and distributed systems preferred.
Python, Java, C++, Rust, UNIX, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Jersey City or McLean or Richmond
HybridFull Time
Exiger
Exiger: Private supply-chain AI software serving corporations, government agencies, and banks with risk and compliance technology.
6+ YOEBachelor's or Master's (or equivalent), 6+ years software/systems engineering with >=4 years in SRE or production/platform reliability, strong Linux/Unix and networking knowledge, experience with SLIs/SLOs, observability, automation, chaos engineering, incident management, and familiarity with AWS and secure/gov environments.
AWS, Codex, Claude, Chaos Monkey, Gremlin, LitmusChaos, Snowflake, Redshift, Apache Iceberg, Go, C, Java, Linux/Unix
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Global financial services and banking institution.
4+ YOE4+ years cloud/platform engineering experience with GCP exposure, Terraform/IaC, CI/CD, observability, DevSecOps practices, incident response, and automation for reliability and resiliency.
GCP, Azure, Terraform, Terraform Enterprise, Log Analytics, Dynatrace, Resource Graph, CI/CD, IAM
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Jersey City or Charlotte or Plano
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Global financial services and banking institution.
4+ YOE4+ years cloud/platform engineering experience with Terraform and GCP; strong IaC, observability, automation, incident response, and DevSecOps skills; ability to define SLIs/SLOs and mentor engineers.
Terraform, Terraform Enterprise, Google Cloud Platform (GCP), Azure, Log Analytics, Dynatrace, Resource Graph, CI/CD
2mo
Save
Mark Applied
Hide
Senior Database Reliability Engineer
San Francisco or New York City or Seattle or Boston or Los Angeles or Chicago or Washington or United States
$145k-$230k/yr HybridFull Time
Scribe
Scribe: Workflow AI software that helps organizations capture, improve, and scale how work gets done.
Deep PostgreSQL and ORM expertise, experience with CDC pipelines (AWS DMS), OpenSearch, Redis, message brokers, observability tools, Python/Go automation, Terraform/IaC, and building reliability/scale for data tiers.
Django, PostgreSQL, Aurora Serverless V2, OpenSearch, Redis, ElastiCache, SQS, RabbitMQ, DMS, S3, Parquet, Snowflake, pganalyze, CloudWatch, Honeycomb, OpenTelemetry, Datadog DBM, pg_stat_statements, Kafka, Python, Go, Terraform, Debezium, Fivetran, Airbyte, pgbouncer, RDS Proxy, Snowpipe, BigQuery, Redshift, SQLAlchemy, ActiveRecord
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yr HybridFull Time
Radix Health
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
AI, HIPAA, PHI, SOC 2, 401(k)
1mo
Save
Mark Applied
Hide
New Technology Introduction Automation Controls Engineer , Global Central Reliability Team
Boston or Austin or Nashville or Bellevue or Arlington or New York
$69k-$115k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
3+ YOEBachelor's in engineering, 3+ years PLC programming and automation controls engineering experience, experience with complex control system design, cross-functional program coordination, and familiarity with industrial control standards and safety.
Allen-Bradley, Siemens, PLC, HMI, SCADA, EtherNet/IP, PROFINET, Modbus, VFDs, servo drives, industrial PC
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Pragma
New York, New York, United States
$175k-$230k/yr HybridFull Time
MarketAxess
MarketAxessNASDAQ: MKTX: Public financial-technology operating an electronic fixed-income trading platform for institutional investors and broker-dealers.
Experience with Java/Python, shell scripting, Linux, SQL, CI/CD (Jenkins), FIX, containers, AWS, monitoring, SRE practices, automation, and strong communication; Bachelor\u0002s/Master\u0002s in CS/Engineering or related field.
Java, Python, bash, ksh, JVM, Linux, SQL, Jenkins, FIX, AWS, GitHub Copilot
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute
San Francisco or New York or Austin or Seattle
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experience owning large GPU/compute fleets, automation of repair/deployment pipelines, firmware/BMC/Redfish familiarity, incident response and paging, observability and metrics tooling, and proficiency with production automation.
Redfish, BMC, IPMI, Temporal, Cadence, Prometheus, Grafana, Go, Python, Kubernetes, Claude Code, Cursor, LLM APIs, MCP servers
4w
Save
Mark Applied
Hide
Senior Lead Platform Reliability Engineer
Charlotte or Irving or Chandler or West Des Moines or Iselin
$159k-$305k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Multinational financial services providing banking and investment products.
7+ YOE7+ years systems engineering or architecture, 5+ years supporting enterprise production environments, deep expertise in one infrastructure domain, SRE practices, automation and troubleshooting across domains.
Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, Python, Bash, PowerShell, Git, Ansible, Terraform, CI/CD
2mo
Save
Mark Applied
Hide
Senior Engineer, Site Reliability, IP
Bethpage, New York, United States
$100k-$165k/yr OnsiteFull Time
Optimum Communications
Optimum CommunicationsNYSE: OPTU: US broadband communications and video services provider.
6+ YOEBachelor's in related field, 6+ years IP networking experience (4+ in reliability), expert routing protocol knowledge, automation (Python, Ansible, Terraform), observability (Grafana, Prometheus, ELK), and leadership/documentation skills.
BGP, OSPF, IS-IS, MPLS, TCP, UDP, ICMP, Python, Ansible, Terraform, Grafana, Prometheus, ELK, SNMP, CI/CD, GitOps, CNI, service mesh, AI Tools
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Private software providing an internal-tools development platform for business and enterprise teams.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Chicago or New York City or Denver
$159k-$172k/yr HybridFull Time
Grubhub
Grubhub: U.S. food-ordering and delivery marketplace connecting diners with local restaurants, merchants, and convenience retailers.
Expertise in infrastructure automation, IaC (Terraform/Pulumi), CI/CD (Jenkins/Spinnaker), containerization (Docker), Linux systems, Python/Bash scripting, AWS and GCP, observability (Vector, Datadog, Splunk), and security for payment environments.
Terraform, Pulumi, Jenkins, Spinnaker, Docker, Amazon EC2, AMI, Linux, Python, Bash, AWS, GCP, Vector, Datadog, Splunk, Ansible, Kubernetes, EKS, GKE
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Cloud, Observability & Automation)
Jersey City, New Jersey, United States
HybridFull Time
DTCC
DTCC: Global post-trade market infrastructure for the financial services industry.
8+ YOEBachelor's degree in computer science, engineering, or equivalent experience; 8+ years in SRE, production engineering, DevOps, or application support; AWS, Python, Java, Go, Linux, monitoring, incident management, and distributed systems expertise.
AWS, Splunk, Grafana, Dynatrace, ITSI, Python, Java, Amazon Q, Kiro, Go, Linux/Unix
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Messaging Services
Secaucus or New York or New Jersey or United States
$140k-$150k/yr RemoteFull Time
National Basketball Association
National Basketball Association: Nonprofit professional basketball league operating five leagues for basketball fans worldwide.
10+ YOEBachelor's in computer science or related, 10+ years in enterprise messaging/infrastructure/reliability, deep Exchange Online/M365 and email security experience, PowerShell and Microsoft Graph automation, Slack/Teams support, incident response and executive support.
Microsoft Exchange Online (M365), Microsoft Outlook, SMTP, Proofpoint, PowerShell, Microsoft Graph, Slack, Microsoft Teams, DMARC, DKIM, SPF, ARC
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
New York City, New York, United States
OnsiteContract
Bahwan CyberTek
Bahwan CyberTek: Global digital transformation providing predictive analytics, AI, and digital supply chain solutions to enterprise clients.
8+ YOERequires 8–12 years of experience with SRE/DevOps, observability, automation, cloud platforms, ETL/ELT, CI/CD, reliability engineering, and large datasets; Python, PowerShell, and Bash skills required.
Python, PowerShell, Bash, Azure DataBricks, Unity Catalog, AWS S3, AWS RDS, ETL, ELT, CI/CD
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Release Engineering
New York, New York, United States
$208k-$274k/yr HybridFull Time
Plaid
Plaid: Fintech data network connecting consumers’ financial accounts to apps and services.
8+ YOE8+ years in backend/SRE/platform engineering; experience designing SLO/SLI programs, progressive delivery, canary rollouts, metric-gated analysis, and automated rollback; proficiency in Go or similar; familiarity with Kubernetes, Prometheus, ArgoCD; strong leadership and incident response skills.
Go, Kubernetes, Prometheus, ArgoCD
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer (FedRAMP)
Bellevue or Chicago or New York City or San Francisco or Washington
$174k-$267k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Identity management and access control software provider.
8+ YOE8+ years operations experience in cloud and Linux, strong networking and web server knowledge, proficiency with Terraform/Chef and scripting (Bash, Python, Go), experience with automation tools and on-call duty.
AWS, Terraform, Chef, Ansible, Puppet, Apache httpd, nginx, Apache Tomcat, Bash, Python, Golang, git, gdb, strace, ltrace, tcpdump, Wireshark, Docker, Kubernetes
2mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer (SRE)
Chicago or New York City
$175k-$220k/yr HybridFull Time
Optimal Market Technologies
Optimal Market Technologies: Private FINRA-registered broker-dealer providing options execution, ATS, routing, and algorithms to retail brokers and institutional trading firms.
Hands-on Linux and network administration, strong scripting/automation, Infrastructure-as-Code experience, production support and incident response ownership, staff management/mentoring, familiarity with C++, Python, SQL, Azure, and real-time trading systems.
C++, Python, SQL, Linux, CentOS7, RHEL9, Microsoft Azure, Microsoft Azure Virtual Desktop (AVD), Claude Code, FIX Protocol, PostgreSQL, AERON
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Axon 911
New York, New York, United States
$141k-$217k/yr HybridFull Time
Carbyne
Carbyne: A building devices and cloud software to improve public safety and emergency response.
5+ YOE5+ years SRE/platform experience; strong observability (Datadog) and Kubernetes experience; AWS familiarity; experience with messaging (RabbitMQ or Kafka); strong monitoring, alerting, deployment, and automation skills.
Datadog, Kubernetes, AWS, RabbitMQ, Kafka