145 infrastructure reliability engineer jobs at 98 companies in Secaucus, NJ

3w
Save
Mark Applied
Hide
Senior Site Reliability and Infrastructure Engineer
New York City, New York, United States
$160k-$220k/yr HybridFull Time
Treeswift
Treeswift: Robotics and physical-AI helping electric utilities’ field crews get more work done.
7+ YOE7+ years experience in observability, SRE, infrastructure or DevOps; hands-on Terraform, Kubernetes, Linux, CI/CD; experience with Airflow-style pipelines and cloud services; strong debugging and communication skills.
Apache Airflow, Astronomer, AWS, Kubernetes, S3, SQS, Lambda, Step Functions, ECS, ECR, Astronomer CLI, Terraform, DuploCloud, Linux
1mo
Save
Mark Applied
Hide
Site Reliability / Infrastructure Engineer
New York City, New York, United States
$180k-$275k/yr OnsiteFull Time
Medal
Medal: Private AI research lab building action and world models for virtual and physical environments.
Experienced with Terraform, Elasticsearch, GCP/Kubernetes, relational database scaling (MySQL/Postgres), incident response, GitHub Actions; strong communication and startup experience preferred.
Terraform, Elasticsearch, GCP, Kubernetes, VPC, IAM, Cloud Logging, MySQL, Postgres, GitHub Actions, CircleCI, Salt, Redis, RabbitMQ, Electron, React, Redux, Styled Components, C#, C++, Swift, Kotlin, Java
1mo
Save
Mark Applied
Hide
Reliability Engineer, R&D
Austin or New York City or San Francisco or Seattle
$203k-$232k/yr OnsiteFull Time
Fluidstack
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experience in reliability engineering for infrastructure or complex hardware, building availability/RAM models, leading cross-discipline FMEAs, and mining field failure data.
2d
Save
Mark Applied
Hide
Customer Reliability Engineer, Infrastructure
United States or Austin or New York City or Boston or San Francisco
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Private software providing managed Apache Airflow data orchestration for enterprise data teams.
5+ YOERequires 5 years of experience with complex cloud infrastructure, 3 years with Kubernetes, production distributed systems, Linux, monitoring, troubleshooting, customer support, DevOps or CI/CD, and Python scripting.
Astro, Apache Airflow, Kubernetes, AWS, GCP, Azure, Linux, Python
2mo
Save
Mark Applied
Hide
Site Reliability Engineer, Tech Infrastructure - USDS
New York, New York, United States
$137k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture LLC
TikTok USDS Joint Venture LLC: Ensuring U.S. data security and content integrity for TikTok.
3+ YOE3+ years SRE/systems engineering experience, bachelor\u0002s in CS or related, proficiency in Python/Go/Java/Shell, Linux, cloud and distributed systems, monitoring and incident management.
Python, Go, Java, Shell, Linux, Docker, Kubernetes, Prometheus, Grafana
1w
Save
Mark Applied
Hide
Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere)
San Francisco or Denver or Austin or Jacksonville or Bridgeport or Seattle or Boston or New York City
$126k-$205k/yr RemoteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Global cybersecurity platform providing network, cloud, and AI-driven security solutions.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
Go, Java, Python, Rust, AWS, GCP, Kubernetes, Linux, Terraform, Cursor, Claude
2w
Save
Mark Applied
Hide
Platform / Site Reliability Engineer
New York City, New York, United States
OnsiteFull Time
Sunset
Sunset: Private startup wind-down service helping founders close companies through legal, tax, and operational work.
Production cloud infrastructure and reliability experience across multiple services, strong software engineering skills, infrastructure and application coding, incident leadership, recovery expertise, and AI engineering tool proficiency.
AWS, Terraform, CI/CD, SOC 2
1w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Brooklyn or New York City or Los Angeles or Santa Monica or United States
$230k-$260k/yr HybridFull Time
Radix Health
Radix Health: Healthcare technology helping providers achieve fair reimbursement through integrated IDR software, data, and AI.
8+ YOE8+ years in SRE, infrastructure, platform engineering, or large-scale production systems; expertise in cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, observability, automation, and incident management.
AI, HIPAA, PHI, SOC 2, 401(k)
2w
Save
Mark Applied
Hide
AI Infrastructure Engineer
New York City, New York, United States
OnsiteFull Time
Palona AI
Palona AI: AI platform helping restaurants capture demand, convert revenue, and manage operations through voice, text, and visual agents.
3+ YOERequires 3+ years in a relevant technical domain, production distributed-systems experience, cloud expertise, infrastructure automation, software development, incident resolution, and knowledge of security, scalability, and reliability.
Python, Docker, AWS, Azure, ECS, Lambda, API Gateway, OpenTofu, Terraform, Datadog, CI/CD
3mo
Save
Mark Applied
Hide
Lead Infrastructure Engineer
New York, New York, United States
$220k-$300k/yr RemoteFull Time
Lakeview Loan Servicing
Lakeview Loan Servicing: Private mortgage servicer and lender that manages residential loans and servicing rights for U.S. homeowners.
5+ YOE2+ Mgmt5+ years building and operating production infrastructure; strong cloud and DevOps experience; hands-on leadership; expertise in IaC and platform reliability.
AWS, GCP, Azure, Terraform, Pulumi, CloudFormation, Bicep, Docker, Kubernetes, ECS, CI/CD, Git, Python, Go, TypeScript
2mo
Save
Mark Applied
Hide
Staff Software Engineer, Infrastructure
Canada or United States or Seattle or Paris or New York City
$238k-$382k/yr RemoteFull Time
Docker
Docker: Privately held container application platform helping developers build, share, and run applications.
8+ YOE8+ years backend/infrastructure engineering; strong Go; experience with Kubernetes/EKS, cloud platforms, networking, and reliability; Bachelor's or equivalent; experience leading cross-team technical initiatives and strong written/verbal communication.
Go, Terraform, Argo CD, EKS, Envoy Gateway, Grafana Cloud, OpenTelemetry, Prometheus, Grafana, GitHub Actions, Kubernetes, Linux, GitOps, Covey Scout
1d
Save
Mark Applied
Hide
Senior Data Infrastructure Engineer
San Francisco or Paris or Seattle or Madrid or London or Berlin or New York City or Sydney or Mexico City
$150k-$200k/yr OnsiteFull Time
Aircall
Aircall: Aircall is a private B2B SaaS providing AI-powered cloud communications software for business sales and support teams.
4+ YOERequires 4+ years in data engineering or infrastructure, strong Python and SQL, Airflow-scale operations, Spark, AWS, Terraform, Kubernetes, Docker, CI/CD, GitOps, reliability ownership, and AI coding tool experience.
Python, SQL, Apache Iceberg, Amazon S3, Flink, Kafka, Amazon MSK, dbt, Apache Kyuubi, Amazon EKS, Amazon Redshift, Rudderstack, Fivetran, DMS, Airflow, Amazon ECS, Lake Formation, StrongDM, SSO, Monte Carlo, Terraform, GitLab CI, GitOps, AWS, Apache Spark, Dagster, Prefect, Amazon IAM, AWS Glue, Amazon Athena, Kubernetes, Docker, Delta Lake, Hudi, Kinesis, Unity Catalog, Claude Code, Cursor
4w
Save
Mark Applied
Hide
Senior Lead Platform Reliability Engineer
Charlotte or Irving or Chandler or West Des Moines or Iselin
$159k-$305k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Multinational financial services providing banking and investment products.
7+ YOE7+ years systems engineering or architecture, 5+ years supporting enterprise production environments, deep expertise in one infrastructure domain, SRE practices, automation and troubleshooting across domains.
Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, Python, Bash, PowerShell, Git, Ansible, Terraform, CI/CD
2w
Save
Mark Applied
Hide
Engineering Lead - Infrastructure and Reliability
New York City, New York, United States
$210k-$237k/yr HybridFull Time
Narmi
Narmi: Fintech providing digital banking, account opening, payments, and fraud-prevention software to community banks and credit unions.
6+ YOE6+ Mgmt6+ years in infrastructure, DevOps, or SRE roles, including engineering leadership; AWS, networking, security, Python, Bash, Linux, system architecture, debugging, and automation skills required.
Amazon Web Services (AWS), Terraform, Postgres, Linux, Python, Bash, ECS, CI/CD, Narmi's Observability platform
1mo
Save
Mark Applied
Hide
Systems Reliability Engineer (SRE)
San Francisco or New York City
$150k-$170k/yr OnsiteFull Time
Claryo
Claryo: AI-powered warehouse intelligence serving logistics operators with AI agents for monitoring, forecasting, and automation.
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Linux, Kubernetes, GCP, AWS, Azure, Prometheus, Grafana, OpenTelemetry, Kafka, RTSP, WebRTC
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco or New York City
$164k-$306k/yr HybridFull Time
Retool
Retool: Private software providing an internal-tools development platform for business and enterprise teams.
Experience operating production infrastructure (AWS), Kubernetes, Terraform, Postgres; programming in Go/Python/TypeScript/Java/Ruby; building observability and automation for customer-facing SaaS systems.
Kubernetes, Helm, Docker Compose, Terraform, AWS, Postgres, Go, Python, TypeScript, Java, Ruby
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York, New York, United States
$178k-$258k/yr RemoteFull Time
Adobe
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
Requires a computer science bachelor's degree or equivalent experience, Python, production ML inference, AWS cloud infrastructure, Kubernetes, vulnerability management, distributed-systems debugging, and on-call participation.
Python, PHP, Node.js, Ruby, SageMaker, OpenAI, Bedrock, EC2 Auto Scaling Groups, Kubernetes, AWS, Azure, GCP, AMI, LangGraph, LLM gateway, MCP, Aurora PostgreSQL, Memcached, Terraform, Terragrunt, Atlantis, Chef, Ansible, SSM, Docker, bash, Jenkins, Argo CD, New Relic, Splunk, Grafana, Prometheus, Fastly, Datadome, WAF, vLLM, LangSmith, Lambda
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Plano or Chandler or Jersey City
$153k-$192k/yr OnsiteFull Time
Bank of America
Bank of AmericaNYSE: BAC: Global financial services and banking institution.
7+ YOE7+ years Azure SRE/cloud infrastructure experience; strong Terraform, observability, networking, scripting (Python/PowerShell/Bash); SRE practices, SLIs/SLOs, incident response, and enterprise platform design.
Microsoft Azure, Terraform, Terraform Enterprise, Azure Monitor, Azure Log Analytics, Dynatrace, Git, Python, PowerShell, Bash, ServiceNow, Jira, Resource Graph, AKS, ACR, Kubernetes, Azure AI Foundry, OpenAI
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Chicago or New York City or Denver
$159k-$172k/yr HybridFull Time
Grubhub
Grubhub: U.S. food-ordering and delivery marketplace connecting diners with local restaurants, merchants, and convenience retailers.
Expertise in infrastructure automation, IaC (Terraform/Pulumi), CI/CD (Jenkins/Spinnaker), containerization (Docker), Linux systems, Python/Bash scripting, AWS and GCP, observability (Vector, Datadog, Splunk), and security for payment environments.
Terraform, Pulumi, Jenkins, Spinnaker, Docker, Amazon EC2, AMI, Linux, Python, Bash, AWS, GCP, Vector, Datadog, Splunk, Ansible, Kubernetes, EKS, GKE
1mo
Save
Mark Applied
Hide
Sr Software Engineer - Reliability Engineering
North Hills, New York, United States
$122k-$203k/yr HybridFull Time
Cox Automotive
Cox Automotive: Privately held automotive services and software serving dealers, fleets, lenders, automakers, and car shoppers.
5+ YOE5+ years in software/platform/infrastructure engineering, strong coding (Python/Go/Java), AWS and Terraform experience, SRE and observability knowledge, system design and incident response skills.
Python, Go, Java, AWS EC2, AWS RDS, AWS DynamoDB, AWS S3, AWS Aurora, AWS Lambda, AWS VPCs, AWS Athena, Terraform, Docker, Kubernetes, Linux, Windows, New Relic, Splunk, Prometheus, CI/CD pipelines