CrewAI
Posted 1mo ago

Software Engineer, AI Runtime & Platform Services

CrewAI
San Francisco, California, United States
HybridFull Time
Responsibilities
  • building runtime
  • maintaining services
  • improving observability
Requirements
  • Strong Python backend/platform experience building production services
  • FastAPI, Celery, Redis, Pydantic, typed Python
  • Distributed systems
  • Auth/security
  • Observability
  • Testing and CI experience
Technical tools mentioned
PythonFastAPICeleryRedisPydanticOpenTelemetrySentrypytestmypyruffAWS ECSAWS ECRKubernetesHelmRails

Job description

About CrewAI

CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production.

The Role

You'll work on the enterprise runtime layer that turns CrewAI's open-source Crews and Flows into secure, observable, remotely executable production systems. This is the layer between the framework and the platform: APIs, workers, checkpoints, webhooks, auth, deployment behavior, telemetry, and enterprise extensions that make CrewAI run reliably in real customer environments.

You'll partner closely with the open-source, product, and infrastructure teams, but your center of gravity is production execution: making agent workflows resumable, inspectable, authenticated, observable, and safe to operate at scale.

What You'll Do

  • Build and maintain the Python enterprise runtime around CrewAI: FastAPI services, Celery workers, Redis-backed state, execution APIs, and deployment-facing tools.
  • Extend open-source CrewAI behavior for enterprise environments while preserving compatibility with upstream framework changes.
  • Own production execution flows: crew and flow kickoff, status, retries, cancellation, checkpoint restore and fork, chat/session state, and human-in-the-loop resume paths.
  • Build secure integration surfaces: JWT auth, signed webhooks, token refresh, file handling, secret fetching, and workload identity across AWS, GCP, and Azure.
  • Improve observability across distributed execution: OpenTelemetry traces, structured logs, Sentry, event tracking, and debuggability across API, worker, and platform boundaries.
  • Maintain strong test coverage for async/runtime behavior using pytest, mypy, ruff, mocks/fakes, and e2e deployment harnesses.
  • Partner with the Agent Management Platform team on API contracts, versioning, enterprise client behavior, deployment status, and failure reporting.

Requirements

What We're Looking For

  • Strong Python backend/platform engineering experience, especially building production services rather than only libraries.
  • Experience with FastAPI or similar API frameworks, Celery or other job systems, Redis, Pydantic, and typed Python.
  • Good instincts for distributed systems: retries, idempotency, async execution, status tracking, race conditions, and failure recovery.
  • Comfort with auth and security-sensitive systems: JWTs, webhooks, signatures, secrets, IAM/workload identity, and least-privilege thinking.
  • Practical observability experience: tracing, structured logging, metrics, Sentry/OpenTelemetry, and debugging multi-service failures.
  • Ability to work at the boundary between an open-source framework and a hosted enterprise platform without creating brittle coupling.
  • Strong testing habits and comfort with CI, package/version management, and release discipline.

Bonus

  • Experience operating AI/agent runtimes, workflow engines, or distributed task systems.
  • Cloud platform experience with AWS ECS/ECR, Kubernetes, Helm, GCP/Azure identity, or secret managers.
  • Experience with enterprise SaaS constraints: auditability, tenant isolation, customer environments, deployment rollbacks, and supportability.
  • Familiarity with Rails/SaaS platforms is useful, but not required.

About CrewAI

Platform for orchestrating collaborative multi-agent AI systems.

Year founded
2023
Employees
50
Organization type
Private
Latest investment
Raised $18.00M Series A (2024) — led by Insight Partners
Headquarters
US

Similar jobs

Software Engineer roles near San Francisco, California
4h
Save
Mark Applied
Hide
Senior Software Engineer(AI/ML), Trust
Bangalore or San Francisco
₹4500k-₹6500k/yr RemoteFull Time
Airbnb
AirbnbNASDAQ: ABNB: Online marketplace for vacation rentals and travel experiences.
7+ YOE7+ years in backend or platform engineering; strong Python or Java, data structures, algorithms, data engineering, machine learning systems, scalable architecture, testing, and deployment experience.
Python, Java, TensorFlow, PyTorch, Kubernetes, Apache Spark, Apache Airflow, Kubeflow, Apache Kafka, Ray, Apache Hive
5h
Save
Mark Applied
Hide
Senior Software Engineer, Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build distributed, large-scale data processing systems, frameworks, and platforms using big data technologies while collaborating with Apple TV and Video teams.
6h
Save
Mark Applied
Hide
Senior Software Engineer, Hyperscale Build Environments and Tools
Santa Clara, California, United States
$180k-$270k/yr OnsiteFull Time
Pure Storage
Pure StorageNYSE: PSTG: Provides all-flash enterprise data storage and management solutions.
5+ YOE5+ years of software engineering experience in infrastructure, developer productivity, build systems, or platform engineering, with C/C++, Linux, Docker, CI, and modern build-system expertise.
C, C++, Make, CMake, Linux, Docker
6h
Save
Mark Applied
Hide
Software Engineer, Linux Kernel / Android
San Jose or Los Angeles or Bellevue
$180k-$240k/yr OnsiteFull Time
Rivet Industries
Rivet Industries: Building integrated task systems for frontline industrial and defense operators.
3+ YOERequires 3+ years developing Android/Linux system software, strong Linux kernel and C/C++ skills, AOSP, HALs, drivers, hardware bring-up, embedded builds, debugging, security mechanisms, and U.S. Person status.
Android, Linux, Linux kernel, AOSP, HALs, C, C++, USB, MIPI, I2C, UART, GPIO, PCIe, U-Boot, Android Bootloader, Yocto, Buildroot, Bazel, Soong, OTA, TPM, AR/XR, NPU, DSP
6h
Save
Mark Applied
Hide
Principal Staff Software Engineer, Systems Infrastructure
Mountain View, California, United States
$226k-$369k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
10+ YOEBA/BS or equivalent practical experience, 10+ years in software or reliability engineering, 5+ years in technical leadership, distributed systems expertise, and experience defining reliability standards across teams.
Java, Go, C++, Python, LLM, SLO, SLI
6h
Save
Mark Applied
Hide
Software Engineer
Bengaluru or San Francisco or Seattle or Ireland
HybridFull Time
DocuSign
DocuSignNASDAQ: DOCU: Provides electronic signature and agreement management software solutions.
5+ YOERequires 5+ years building production backend or platform services, service design ownership, backend programming, databases, APIs, testing, version control, CI/CD, and distributed systems knowledge.
Java, C#, Go, TypeScript, Node.js, SQL, NoSQL, Kubernetes, Docker, Spark, Flink, Kafka, RESTful, CI/CD
8h
Save
Mark Applied
Hide
Exceptional Software Engineer
Redwood City, California, United States
$180k-$400k/yr OnsiteFull Time
Dyna Robotics
Dyna Robotics: Develops general-purpose robots powered by proprietary embodied AI models.
Exceptional software engineering ability, high tolerance for ambiguity, resilience in changing environments, and clear communication. Robotics experience is welcome but not required.
8h
Save
Mark Applied
Hide
Staff+ Software Engineer, Product Sandboxing
San Francisco or New York City or Seattle or California
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOERequires 8+ years building scalable distributed systems, strong service-oriented architecture, networking, and systems design expertise, proficiency in Python, Go, or Rust, and experience with cloud infrastructure and Kubernetes.
Python, Go, Rust, GCP, AWS, Azure, Kubernetes