Vibecoderz
Posted 10mo ago

DevOps Engineer (Founding Team)

Vibecoderz
Hyderābād or San Francisco
$24k-$32k/yrHybridFull Time
Responsibilities
  • Own CI/CD
  • Define IaC
  • Set observability
Requirements
  • 10+ years in DevOps/SRE with CI/CD, IaC, and cloud security
  • Strong observability and cost optimization experience
Technical tools mentioned
TerraformGCP (Cloud Run, Firestore, Pub/Sub, VPC)GitHub ActionsOpenTelemetryPrometheusGrafanaCloud TraceIAMSecrets ManagerRedisNeo4jBrowser-Use scaling infra

Job description

Role Overview

You will be the guardian of infrastructure and velocity. Every AI artifact, every mini-app, every TutorAgent interaction depends on the reliability of the pipelines and cloud environment you build. At Vibecoderz, DevOps isn’t a support role — it’s the backbone of product velocity and reliability.

As the founding DevOps engineer, you will design our CI/CD pipelines, observability stack, infra-as-code, and security posture from day zero. You’ll ensure that when a developer ships code, it’s live in production safely, quickly, and traceably. You’ll architect the serverless-first cloud strategy on GCP, balancing performance, cost, and scale as we grow from MVP to 100K+ MAUs.

You’ll use Linear (execution), Notion (runbooks/docs), and GitHub (actions & infra code) to make every workflow reproducible, transparent, and automated.

Key Responsibilities

  1. CI/CD Pipeline Ownership

    • Design and maintain GitHub Actions workflows for FE, BE, AI, and agent services.

    • Automate build, test, and deploy to Cloud Run with zero-downtime releases.

  2. Infrastructure as Code (IaC)

    • Implement Terraform scripts for GCP (Cloud Run, Firestore, Pub/Sub, VPCs).

    • Maintain environment parity (dev, staging, prod).

  3. Observability & Monitoring

    • Set up OpenTelemetry tracing for multi-agent workflows.

    • Configure dashboards (Cloud Trace, Grafana) for latency, errors, and throughput.

  4. Cost & Resource Optimization

    • Track infra costs, optimize workloads, and enforce scaling policies.

    • Benchmark agent workloads across Gemini Flash vs. Pro vs. custom models.

  5. Cloud Security & Compliance

    • Enforce IAM best practices, firewall rules, and secret management.

    • Build guardrails for prompt injection and unsafe agent actions at the infra level.

  6. Release Management

    • Define release pipelines with feature flags, rollbacks, and canary deploys.

    • Ensure smooth collaboration between PM, engineers, and QA.

  7. Disaster Recovery & Resilience

    • Build automated backup and recovery strategies for Firestore + Neo4j + Redis.

    • Design failover strategies for critical agent services.

  8. Agent Infrastructure Support

    • Support Browser-Use scaling for the Vibe Browser.

    • Manage GPU/TPU allocations for Gemini/Vertex pipelines if required.

  9. Collaboration & Enablement

    • Write runbooks and incident playbooks in Notion.

    • Train engineering team to self-serve common workflows.

  10. Problem Solving

    • Debug infra bottlenecks, trace latency across services, and enforce SLAs.

Success Metrics

90 Days (Probation):

  • CI/CD pipeline live for FE + BE services.

  • Terraform-based infra deployed and reproducible.

  • OpenTelemetry traces visible for at least 2 core user flows.

12 Months:

  • 99.9% uptime across production workloads.

  • <200ms latency for API responses across multi-agent workflows.

  • Fully automated deployments with rollback & feature flag system.

  • Disaster recovery tested with <5 min RTO (Recovery Time Objective).

Must-Haves

  • 10+ years in DevOps/SRE roles for high-scale products.

  • Mastery of CI/CD, Terraform, GCP services (Cloud Run, Pub/Sub, Firestore).

  • Proven experience with observability stacks (OpenTelemetry, Prometheus, Grafana).

  • Deep knowledge of cloud security, IAM, and infra cost management.

  • Background in scaling infra for developer or AI products.

Nice-to-Haves

  • Experience with Vertex AI/ML infra and GPU/TPU scaling.

  • Prior work on real-time, multi-agent systems.

  • Contributions to open-source DevOps tooling.

  • Startup/founding engineer experience.

Tech Stack Visibility

  • Infra: Terraform, GCP (Cloud Run, Pub/Sub, Firestore, VPC)

  • CI/CD: GitHub Actions

  • Observability: OpenTelemetry, Cloud Trace, Grafana

  • Security: IAM, Secrets Manager, GCP Firewall

  • Other: Redis, Neo4j, Browser-Use scaling infra

Assessment

Objective: Validate ability to design and operate production-grade infra for Vibecoderz.

Challenge (Candidate PoC):

  1. CI/CD Setup

    • Create a GitHub Actions workflow to:

      • Run unit tests for FE (Next.js) + BE (FastAPI).

      • Deploy BE service to Cloud Run on merge to main.

  2. IaC

    • Write Terraform scripts to provision:

      • Cloud Run service

      • Firestore DB

      • Pub/Sub topic for agent comms

  3. Observability

    • Add OpenTelemetry traces for one workflow (Text → Course).

    • Export traces to Cloud Trace and provide a screenshot of latency breakdown.

  4. Security

    • Configure IAM policy with least-privilege roles.

    • Add secrets management (e.g., API keys) to the workflow.

Deliverables:

  • GitHub repo with workflows + Terraform configs.

  • Cloud Run URL for deployed BE service.

  • Tracing screenshot with latency insights.

  • Short README explaining infra choices + tradeoffs.

Evaluation Criteria:

  • CI/CD Workflow Robustness (25%)

  • IaC Quality & Reproducibility (25%)

  • Observability & Monitoring Depth (20%)

  • Security & IAM Best Practices (15%)

  • Documentation & Clarity (15%)

About Vibecoderz

A building AI infrastructure and tooling to power scalable AI products.

Similar jobs

DevOps Engineer roles near Hyderābād, Telangana
3h
Save
Mark Applied
Hide
DevOps Engineer
Hyderabad, Telangana, India
OnsiteFull Time
Gradera
Gradera: AI-native platform and services for orchestrated enterprise business transformation.
5+ YOEDevOps, SRE, or infrastructure engineering experience; expertise in cloud platforms, Docker, Kubernetes, infrastructure as code, CI/CD, observability, security, and distributed systems. 5+ years preferred.
AWS, GCP, Azure, Kubernetes, Docker, Terraform, CloudFormation, GitHub Actions, GitLab CI, Jenkins
10h
Save
Mark Applied
Hide
DevOps Engineer
Hyderabad, Telangana, India
OnsiteFull Time
Accenture
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
5+ YOERequires 5+ years in DevOps Architecture, 15 years of full-time education, and expertise in CI/CD, cloud infrastructure, container orchestration, security, Python, AWS Architecture, and AWS CloudFormation.
Python, AWS, AWS CloudFormation, CI/CD
20h
Save
Mark Applied
Hide
DevOps Engineer
Hyderabad or Visakhapatnam or Hyderabad
HybridFull Time
XTGlobal
XTGlobalBSE: XTGL: Provider of IT consulting, software development, and BPO services.
8+ YOERequires 8+ years in DevOps or related engineering, with Kubernetes, GitOps, Terraform, Azure, GCP, CI/CD, cloud architecture, troubleshooting, solution design, and technical leadership experience.
Kubernetes, GitOps, Terraform, Microsoft Azure, Google Cloud Platform (GCP), CI/CD, Docker, Infrastructure as Code (IaC), Agile/Scrum
1d
Save
Mark Applied
Hide
Manager_DevOps
Hyderabad or Bengaluru
HybridFull Time
NTT DATA Business Solutions
NTT DATA Business Solutions: Global SAP partner providing IT consulting and managed services.
10+ YOERequires 10+ years of IT experience and 7+ years of hands-on DevOps experience with AWS, Kubernetes, Docker Enterprise, CI/CD, Terraform, infrastructure automation, platform administration, and monitoring.
AWS, Kubernetes, Docker Enterprise, Terraform, GitHub Actions, GitHub CI/CD, EC2, IAM, S3, CloudWatch, Lambda, Linux, Unix, Windows Server, WebLogic, WebSphere, Tomcat, Nginx
1d
Save
Mark Applied
Hide
Lead II - DevOps Engineering
Hyderabad, Telangana, India
OnsiteFull Time
UST
UST: Global provider of digital transformation and IT services.
7+ YOERequires 7+ years with Azure services, 5+ years with Terraform, Azure DevOps CI/CD, security scanning integrations, scripting, and identity platforms; regulated-industry experience and Azure certifications preferred.
Microsoft Azure DevOps, CI/CD, Docker, Azure Kubernetes Service (AKS), Azure API Management, Terraform, Wiz, Checkmarx, Okta, Python, Bash, PowerShell, Microsoft Azure
1d
Save
Mark Applied
Hide
Sr. Staff DevOps Engineer
Hyderabad, Telangana, India
OnsiteFull Time
GE Vernova
GE VernovaNYSE: GEV: Designs and services technologies for global power generation and electrification.
Extensive DevOps experience with CI/CD, Kubernetes, Docker, Infrastructure as Code, cloud platforms, monitoring, observability, DevSecOps, and enterprise modernization; bachelor's or master's preferred.
Jenkins, Kubernetes, Docker, Helm, Terraform, Ansible, Git
2d
Save
Mark Applied
Hide
DevOps IRC302483
Hyderabad, Telangana, India
HybridFull Time
GlobalLogic
GlobalLogic: Digital product engineering and software development services provider.
5+ YOERequires 5+ years of Kubernetes experience, strong AWS, Helm, Jenkins, Docker, Python 3.10+, REST API, incident response, and PostgreSQL skills; EKS experience preferred.
Kubernetes, Amazon EKS, Jenkins, Amazon EC2, Auto Scaling Group (ASG), AWS ALB/NLB, AWS CloudTrail, Amazon CloudWatch, AWS IAM, Amazon Route 53, AWS STS, Helm, Docker, Amazon ECR, YAML, Python, FastAPI, pytest, black, isort, mypy, pydantic, httpx, boto3, Jira, Confluence, LDAP/AD, PostgreSQL, LangChain, LangGraph, Anthropic, OpenAI, Microsoft Excel
2d
Save
Mark Applied
Hide
Dev Ops Manager - Azure
Hyderabad, Telangana, India
OnsiteFull Time
Blend360
Blend360: Provides data science, AI, and marketing consulting services.
Requires hands-on production experience with containerized microservices, Azure Container Apps or Databricks Apps, Docker, enterprise CI/CD integration, monitoring, secrets management, RBAC, Git workflows, and production AI or data platforms.
Azure Container Apps, Databricks Apps, Docker, Microsoft Azure, Azure DevOps, Databricks, Git, CI/CD, RBAC