Helios Intelligence Platforms
Posted 5d ago

Staff/Senior Distributed Systems Engineer

Helios Intelligence Platforms
New York City, New York, United States
$185k-$235k/yrOnsiteFull Time
Responsibilities
  • operating platforms
  • managing deployments
  • optimizing queues
Requirements
  • Expertise in distributed execution
  • Scheduling
  • Storage
  • Reliability
  • Cloud infrastructure, CI/CD
  • Observability
  • Security, and performance optimization for mission-critical systems
Technical tools mentioned
KafkaRedpandaPulsarSQSPub/SubRabbitMQTemporalPostgreSQLRedisElasticsearchTypesenseAWSGCPAzureTerraformOpenTelemetryDatadogMicrosoft

Job description

Staff/Senior Distributed Systems Engineer

New York City | Full-time | On-site in SoHo, five days per week

Reports to the CTO and works directly with the founding team.

ABOUT HELIOS

Helios is building a new kind of company to solve America’s hardest problems, starting with the government interaction layer.

Government shapes every consequential market, but the infrastructure connecting public institutions and private organizations remains fragmented, manual, and difficult to navigate. Helios is rebuilding that layer.

Our core platform, Proxi, gives organizations the intelligence they need to understand what government is doing, why it matters, and what to do next. From that foundation, we design and deploy secure, mission-specific systems for government agencies, enterprises, and institutions operating in complex and highly regulated environments.

We bring together frontier AI, deep public-sector expertise, and forward-deployed execution. Our team includes leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft. We are backed by leading institutional investors and trusted by organizations working on high-stakes problems across government and industry.

MISSION

As a distributed systems engineer at Helios you will expand and operate the distributed execution, storage, scheduling, and reliability primitives that support data ingestion, document processing, search indexing, model inference, agents, and continuously running research workflows.

The primary responsibility centers around the orchestration of critical infrastructure to meet mission-critical performance and reliability guarantees. Our platform runs a multitude of network crawling, compute grids, memory-heavy document processing, OCR, inference, and latency-sensitive jobs, all of which you should be familiar with. Preferred areas of expertise include:

  • Queue-, log-, workflow-, and actor-based distributed execution systems, including Kafka, Redpanda, Pulsar, SQS, Pub/Sub, RabbitMQ, Temporal, and equivalent technologies.

  • Establishment of at-least-once delivery with effectively once-only business outcomes through transactional outbox and inbox patterns, sagas, reconciliation, dead-letter handling, replay and historical backfill procedures.

  • CPU-, memory-, disk-, network-, and GPU-aware scheduling, including workload classification, priority allocation, starvation prevention, tenant and dependency concurrency limits, placement constraints, resource quotas, and noisy-neighbor isolation.

  • PostgreSQL, object storage, Redis or equivalent caches, search indices (Elasticsearch or Typesense), vector stores, change-data-capture, and graph-storage systems.

  • Application of distributed coordination and concurrency controls, including leader election, locks, leases, fencing tokens, optimistic concurrency, conflict resolution, and partition recovery.

  • Provisioning and operation of AWS, GCP, or Azure infrastructure through terraform or similar IaC, vulnerability scanning, CI/CD and deployment methods.

  • Definition and administration of SLO/Is, error budgets, release controls, metrics, OpenTelemetry, Datadog, fault injection and load/failure testing.

  • Evaluation of system performance and cost via e2e latency decomposition, resource and dependency profiling.

  • Enforcement of platform security and tenant isolation policies.

The key objective of this work is to support the real-time access and availability of our entire data corpus as well as supporting the continued construction of the Helios Rapid Ontology System (H.R.O.S.), our long horizon memory data plane.

KEY RESPONSIBILITIES

  • Own the operation and architecture of Proxi’s distributed execution platform, covering ingestion, document processing, search index performance and long-running research jobs.

  • Manage the deployment of our GovCloud and Air-Gapped resources for sensitive environments.

  • Manage the orchestration of long-running agent research tasks including scheduling, lease management and retention.

  • Queue optimization and cross-cloud information pipeline scalability.

  • Build out dedicated resource-aware autoscaling architecture for fast search and document processing resources.

  • Manage CI/CD and compliance operations across the entire Helios platform.

  • Support global forward embedded customer infrastructure efforts.

WORKING AT HELIOS

This is a full-time, in-person role based in our SoHo office in New York City. Team members are expected to work from the office five days per week.

Certain customer engagements may require background checks, security reviews, access approvals, or eligibility for a U.S. government security clearance. Some projects may be subject to U.S. citizenship or other customer-specific access requirements.

We're not looking for passengers; we want driven innovators with a hunger to build from the ground up - obsessed with pushing the boundaries of natural language understanding, document processing, and personalized relevance.

Helios is a fast-moving startup with ambitious goals. This is not a conventional 9-to-5 role. We expect flexibility during critical deployment periods, customer incidents, product launches, and other company-critical work. In return, this role offers unusual ownership, direct access to consequential institutions, and the opportunity to build systems that affect how major decisions are made.

HOW WE WORK

  • Own outcomes, not just assigned tasks. Identify what needs to happen and drive it through completion.

  • Move quickly without lowering the standard. Speed and rigor are complementary.

  • Stay close to the mission and the user. The best decisions begin with the real problem.

  • Work across boundaries. Everyone contributes beyond the narrow limits of a job title.

  • Communicate directly. We value clear thinking, honest feedback, and low-ego collaboration.

  • Build for the real world. Our systems must perform in complex, regulated, and high-stakes environments.

EQUAL OPPORTUNITY

Helios is an equal opportunity employer. We evaluate candidates based on their abilities, experience, and potential to contribute to our mission. We do not discriminate on the basis of race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other status protected by applicable law.

About Helios Intelligence Platforms

AI policy-intelligence platform and consulting firm serving government, lobbying, legal, compliance, and enterprise policy teams.

Similar jobs

Distributed Systems Engineer roles near New York City, New York
5d
Save
Mark Applied
Hide
Senior Distributed Systems Engineer
United States or Canada or San Francisco or New York City or Seattle or Ann Arbor
$151k-$191k/yr RemoteFull Time
Censys
Censys: Internet intelligence and cybersecurity software providing threat hunting and attack-surface management to governments and enterprises.
5+ YOERequires 5+ years of software engineering experience building distributed systems, object-oriented programming in Go, cloud provider experience, and familiarity with message queues, databases, AI, and maintainable code.
Go, AWS, Azure, GCP, AWS Kinesis, Google Pub/Sub, Kafka, BigTable, Cloud Spanner, HBase, Cassandra, gRPC, REST, Protobuf, MessagePack, Kubernetes, AI
1mo
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
1mo
Save
Mark Applied
Hide
Distributed Systems Engineer
San Francisco or New York City or Austin or Seattle
$175k-$300k/yr OnsiteFull Time
Fluidstack
Fluidstack: Building and operating civilization-scale data center infrastructure for AI.
Experienced engineer building observability, control planes, and API surfaces for large fleets; production service ownership, pager/incident experience, fluent with AI tooling and distributed systems.
LLM APIs, MCP servers, Claude Code, Cursor, Prometheus, Thanos, VictoriaMetrics, Temporal, Cadence, BMC/Redfish, Go, Python, Postgres, Kubernetes, ZTP, DHCP, DNS
2mo
Save
Mark Applied
Hide
Distributed Systems Engineer SMTS/LMTS
New York City or San Francisco
$149k-$224k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: The #1 AI CRM driving customer success together.
Design and deliver scalable, reliable, secure backend distributed systems within the Security ecosystem; hands-on engineering, cross-functional collaboration, and end-to-end ownership with an AI-first mindset.
Salesforce Core, Data Cloud, Tableau Next, ATF, Agentforce
4mo
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years building distributed systems; ads domain experience; ML model serving; backend APIs and ad tech systems; low-latency serving; collaboration across teams.
ML model serving, real-time inference, APIs, ad servers, bidders, pacing, routing, calibration serving
5mo
Save
Mark Applied
Hide
Member of Technical Staff - Distributed Systems Engineer
New York City or San Francisco or London
OnsiteFull Time, Internship
Reflection AI
Reflection AI: Private AI research lab building open foundation models and agentic software for enterprise, government, and regulated-industry users.
Strong software engineering background; production-grade systems; API/service/platform design for large-scale data/compute; collaborative in a fast-paced startup.
Go, Rust, C++, gRPC, Protobuf, Kubernetes, Docker, APIs
1y
Save
Mark Applied
Hide
Distributed Systems Engineer
New York City, New York, United States
$200k-$325k/yr OnsiteFull Time
Moment
Moment: Private AI software for investment management serving large wealth firms and fintechs.
Experience building distributed systems, scalable data pipelines, and real-time data platforms; proficient in backend software design and systems programming.
Golang, gRPC, PostgreSQL, Kafka, ClickHouse, AWS Lambda
2y
Save
Mark Applied
Hide
Distributed Systems Engineer
New York City or United States
$125k-$250k/yr HybridFull Time
Axiom
Axiom: A software building zero-knowledge infrastructure for developers in crypto and fintech.
Distributed systems engineer with Rust/C++ experience; cloud deployment and security mindset; strong collaboration and communication.
Rust, C++, Docker, Kubernetes, Terraform