Demandbase
Posted 3w ago

DevOps Engineer - AI Runtime Services

Demandbase
Hyderabad, Telangana, India
OnsiteFull Time
Responsibilities
  • building runtime
  • managing reliability
  • controlling costs
Requirements
  • Production infrastructure experience with Kubernetes
  • AWS/GCP
  • Terraform
  • GitOps
  • Strong Python
  • On-call and incident-response experience
  • Observability instrumentation
  • Hands-on LLM system experience
Technical tools mentioned
KubernetesAWSGCPEKSTerraformGitOpsFluxKarpenterPythonLiteLLMDatadog LLM ObservabilityDatadogPrometheusLokiClickHouseGitLab CILangSmithBraintrustArize

Job description

Introduction to Demandbase:

Demandbase is the only pipeline AI platform that empowers GTM teams to automate growth at scale. With a unified view of data, insights, actions, and outcomes, B2B enterprises can seamlessly align and execute their account-based GTM strategies with confidence. Thousands of businesses trust Demandbase to maximize revenue, minimize waste, and consolidate their data and tech stacks – all in one platform.

As a company, we’re as committed to growing careers as we are to building world-class technology. We invest heavily in people, our culture, and the community around us. We have also continuously been recognized as One of The Best Places To Work in the San Francisco Bay Area by Fortune, and One of The 60 Best Companies To Sell For by Selling Power. Our offices are located in San Francisco, New York, Austin, Seattle, India, and the United Kingdom.

About the Role

AI Runtime Services owns the shared runtime that AI at Demandbase runs on — two ways. First, it's the layer that enables and supports the AI features in the Demandbase platform: product teams ship LLM features on top of it without reinventing model access, spend control, safety, and observability. Second, the team builds and supports internal AI solutions — the tooling and services the company uses to work with LLMs day to day. The pillars underneath both: the LLM gateway, cost controls, observability, evals infra, caching, and guardrails.

The surface is broad and the team is early. You won't be handed a narrow slice: expect to move across the gateway one week and the ingest pipeline the next — and across product-facing runtime and internal tooling — and to own what you build in production.

What you'll own

  • The LLM gateway. The single front door to every model provider we use — keys, routing, rate limits, failover, provider quotas. When it's down, every AI feature at Demandbase is down.

  • Reliability, for real. SLOs, on-call, incident response, and the postmortems for the runtime. This is a genuine SRE ownership role, not "build it and let someone else run it."

  • Cost control. Per-team and per-model attribution, budgets, quotas. LLM spend is the kind of line item that quietly triples; your job is to make it legible and then make it smaller.

  • LLM observability. Traces, spans, prompt/response capture, ingest and billing visibility for every AI feature in the company.

  • Evals and experimentation infra. The shared harnesses, datasets, and scoring plumbing product teams use to know whether a prompt or model change actually made things better.

  • Caching and guardrails. Response and semantic caching to cut latency and redundant spend; input/output safety, PII handling, and policy enforcement at the gateway.

What we need

  • Strong production infrastructure background: Kubernetes, AWS, Terraform, GitOps. You've operated systems that people were paged for, and you've been the one paged.

  • Python, at a level where you're comfortable owning services and tooling in it — not just scripting.

  • Real on-call and incident-response experience. You can talk about an incident you ran, what you got wrong, and what you changed afterward.

  • Observability fluency beyond "we have dashboards": you've instrumented systems, chased cardinality and cost in a metrics/logging backend, and built alerts that fire when they should.

  • Hands-on familiarity with how LLM systems actually work. You don't need production ML experience. But you need to have built something with these tools — an agent, a RAG pipeline, an internal tool — and to have used evals or experiments to decide whether it was any good. Tokens, context windows, prompt/response tracing, and why an eval suite is fundamental should all be familiar ground. Candidates who have only read about this are not a fit.

Nice to have

  • An LLM gateway or proxy (LiteLLM or similar) in production.

  • LLM observability tooling: Datadog LLM Observability, LangSmith, Braintrust, Arize, or similar.

  • Eval frameworks and LLM-as-judge scoring in a real workflow, not a demo.

  • FinOps instincts — cost attribution, showback, quota design.

  • Guardrails, PII detection, or content-filtering systems.

  • GCP alongside AWS.

Our stack

AWS - GCP · EKS, Flux/GitOps, Karpenter · Python · LiteLLM · Datadog (LLM observability, APM) · Prometheus, Loki, ClickHouse · Terraform · GitLab CI

Benefits

Our benefits include Group Medical, Personal Accident, and Term Life Insurance for comprehensive protection. Preventive healthcare covers dental, vision, and OPD needs, complemented by strong mental health support. We also provide a fitness benefit, car lease policy, and gratuity for long-term financial well-being.

Our Commitment to Diversity, Equity, and Inclusion at Demandbase

At Demandbase, we believe in creating a workplace culture that values and celebrates diversity in all its forms. We recognize that everyone brings unique experiences, perspectives, and identities to the table, and we are committed to building a community where everyone feels valued, respected, and supported. Discrimination of any kind is not tolerated, and we strive to ensure that every individual has an equal opportunity to succeed and grow, regardless of their gender identity, sexual orientation, disability, race, ethnicity, background, marital status, genetic information, education level, veteran status, national origin, or any other protected status. We do not automatically disqualify applicants with criminal records and will consider each applicant on a case-by-case basis.

We recognize that not all candidates will have every skill or qualification listed in this job description. If you feel you have the level of experience to be successful in the role, we encourage you to apply!

We acknowledge that true diversity and inclusion requires ongoing effort, and we are committed to doing the work required to make our workplace a safe and equitable space for all. Join us in building a community where we can learn from each other, celebrate our differences, and work together.

Unsolicited Submissions

At Demandbase, we value thoughtful partnerships and direct connections with candidates. We’re not accepting unsolicited resumes or outreach from third-party recruiting agencies. Any unsolicited submissions will not be reviewed, and no fees will be paid.

About Demandbase

Provides account-based marketing and sales intelligence software.

Year founded
2006
Employees
750
Organization type
Private
Latest investment
Raised $175.00M Series H (2023) — led by Silver Lake Waterman
Subsidiaries
Headquarters
US

Similar jobs

DevOps Engineer roles near Hyderabad, Telangana
21h
Save
Mark Applied
Hide
Manager_DevOps
Hyderabad or Bengaluru
HybridFull Time
NTT DATA Business Solutions
NTT DATA Business Solutions: Global SAP partner providing IT consulting and managed services.
10+ YOERequires 10+ years of IT experience and 7+ years of hands-on DevOps experience with AWS, Kubernetes, Docker Enterprise, CI/CD, Terraform, infrastructure automation, platform administration, and monitoring.
AWS, Kubernetes, Docker Enterprise, Terraform, GitHub Actions, GitHub CI/CD, EC2, IAM, S3, CloudWatch, Lambda, Linux, Unix, Windows Server, WebLogic, WebSphere, Tomcat, Nginx
1d
Save
Mark Applied
Hide
Lead II - DevOps Engineering
Hyderabad, Telangana, India
OnsiteFull Time
UST
UST: Global provider of digital transformation and IT services.
7+ YOERequires 7+ years with Azure services, 5+ years with Terraform, Azure DevOps CI/CD, security scanning integrations, scripting, and identity platforms; regulated-industry experience and Azure certifications preferred.
Microsoft Azure DevOps, CI/CD, Docker, Azure Kubernetes Service (AKS), Azure API Management, Terraform, Wiz, Checkmarx, Okta, Python, Bash, PowerShell, Microsoft Azure
1d
Save
Mark Applied
Hide
Sr. Staff DevOps Engineer
Hyderabad, Telangana, India
OnsiteFull Time
GE Vernova
GE VernovaNYSE: GEV: Designs and services technologies for global power generation and electrification.
Extensive DevOps experience with CI/CD, Kubernetes, Docker, Infrastructure as Code, cloud platforms, monitoring, observability, DevSecOps, and enterprise modernization; bachelor's or master's preferred.
Jenkins, Kubernetes, Docker, Helm, Terraform, Ansible, Git
1d
Save
Mark Applied
Hide
DevOps Engineer
Hyderabad, Telangana, India
OnsiteFull Time
Accenture
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
3+ YOERequires 3+ years of experience, 15 years of full-time education, and proficiency in SAP S/4HANA Advanced Available to Promise; cloud, container orchestration, security, and CI/CD experience required.
SAP S/4HANA Advanced Available to Promise, Accenture Delivery Methods (ADM), CI/CD
2d
Save
Mark Applied
Hide
DevOps IRC302483
Hyderabad, Telangana, India
HybridFull Time
GlobalLogic
GlobalLogic: Digital product engineering and software development services provider.
5+ YOERequires 5+ years of Kubernetes experience, strong AWS, Helm, Jenkins, Docker, Python 3.10+, REST API, incident response, and PostgreSQL skills; EKS experience preferred.
Kubernetes, Amazon EKS, Jenkins, Amazon EC2, Auto Scaling Group (ASG), AWS ALB/NLB, AWS CloudTrail, Amazon CloudWatch, AWS IAM, Amazon Route 53, AWS STS, Helm, Docker, Amazon ECR, YAML, Python, FastAPI, pytest, black, isort, mypy, pydantic, httpx, boto3, Jira, Confluence, LDAP/AD, PostgreSQL, LangChain, LangGraph, Anthropic, OpenAI, Microsoft Excel
2d
Save
Mark Applied
Hide
Dev Ops Manager - Azure
Hyderabad, Telangana, India
OnsiteFull Time
Blend360
Blend360: Provides data science, AI, and marketing consulting services.
Requires hands-on production experience with containerized microservices, Azure Container Apps or Databricks Apps, Docker, enterprise CI/CD integration, monitoring, secrets management, RBAC, Git workflows, and production AI or data platforms.
Azure Container Apps, Databricks Apps, Docker, Microsoft Azure, Azure DevOps, Databricks, Git, CI/CD, RBAC
2d
Save
Mark Applied
Hide
DevOps Engineer
Hyderabad, Telangana, India
OnsiteFull Time
Accenture
AccentureNYSE: ACN: Global provider of management consulting and technology services.
3+ YOERequires 3+ years in ServiceNow SPM, 15 years of full-time education, and skills in CI/CD, cloud platforms, container orchestration, security, scalable infrastructure, automation, and scripting.
ServiceNow Strategic Portfolio Management (SPM)
2d
Save
Mark Applied
Hide
Principal DevOps Engineer
Hyderabad, Telangana, India
OnsiteFull Time
Blackbaud
BlackbaudNasdaq: BLKB: Provides software for nonprofits and social impact organizations.
8+ YOEBachelor’s degree plus 8+ years of related experience or equivalent; 8+ years managing Linux and RedHat infrastructure, web technologies, Azure, cloud models, automation, scalable systems, and secure production operations.
CI/CD, Linux, RedHat, JavaScript, Java, HTML, AJAX, Microsoft Azure, SaaS, PaaS, IaaS, Perl, PowerShell, Bash, Ansible, Python, Terraform