Kibo
Posted 4mo ago

DevOps Engineer

Kibo
Pune, Maharashtra, India
RemoteFull Time
Responsibilities
  • Manage clusters
  • Troubleshoot issues
  • Build infrastructure
Requirements
  • 8+ years DevOps/Developer Engineer
  • Ownership of production Kubernetes (EKS preferred)
  • Real-time troubleshooting
  • Terraform
  • On-call rotation
Technical tools mentioned
KubernetesEKSTerraformPrometheusGrafana

Job description


About This Role:


We are hiring a hands-on DevOps Engineer to manage and support production-grade cloud infrastructure for Kibo’s commerce platform. This role focuses on Kubernetes (EKS), Terraform, and real-time production troubleshooting in a 24/7 on-call environment.




ABOUT KIBO 




KIBO is a composable digital commerce platform for B2C, D2C, and B2B organizations who want to simplify the complexity in their businesses and deliver modern customer experiences.  KIBO is the only modular, modern commerce platform that supports experiences spanning B2B and B2C Commerce, Order Management, and Subscriptions. Companies like Ace Hardware, Zwilling, Jelly Belly, Nivel, and Honey Birdette trust Kibo to bring simplicity and sophistication to commerce operations and deliver experiences that drive value.   


KIBO's cutting-edge solution is MACH Alliance Certified and has been recognized by Forrester, Gartner, IDC, Internet Retailer, and TrustRadius. KIBO has been named a leader in The Forrester Wave™: Order Management Systems, Q1 2025 and in the IDC MarketScape report “Worldwide Enterprise Headless Digital Commerce Applications 2024 Vendor Assessment”.


By joining KIBO, you will be part of a team of Kibonauts all over the world in a remote-friendly environment. Whether your job is to build, sell, or support KIBO’s commerce solutions, we tackle challenges together with the approach of trust, growth mindset, and customer obsession. If you’re seeking a unique challenge with amazing growth potential, then come work with us!


 



WHAT YOU’LL DO 





  • Manage and operate production-grade Kubernetes clusters (EKS preferred), ensuring high availability and scalability

  • Troubleshoot real-time production issues across distributed systems and microservices

  • Diagnose and resolve issues such as:

    • Pod failures (CrashLoopBackOff, Pending, OOMKilled)

    • Node failures, autoscaling, and resource constraints

    • Networking, ingress, and service connectivity issues



  • Build, maintain, and debug infrastructure using Terraform (modules, remote state, locking, drift handling)

  • Implement and enhance monitoring & alerting systems using Prometheus, Grafana, and related tools

  • Perform root cause analysis (RCA) for incidents and drive permanent fixes to improve system reliability

  • Participate in a 24/7 on-call rotation, owning incidents and resolving them independently

  • Collaborate with engineering teams to improve system performance, resilience, and deployment processes

  • Automate deployments, infrastructure provisioning, and operational workflows to reduce manual effort

  • Ensure adherence to security best practices across infrastructure and deployments 


 



WHAT YOU’LL NEED 







  • 8 + Years of experience as a Developer Engineer, owning and operating production Kubernetes clusters (EKS preferred), including cluster health, scaling, and availability

  • Troubleshoot real-time production issues independently across microservices and distributed systems

  • Debug and resolve critical issues such as:

    • Pods stuck in CrashLoopBackOff, Pending, OOMKilled states

    • Node failures, node pressure, autoscaling issues

    • Service connectivity, ingress, and networking issues



  • Investigate and fix cluster-level issues including scheduling, resource constraints, and misconfigurations

  • Build and maintain infrastructure using Terraform, including:

    • Writing and modifying modules

    • Managing remote state and locking

    • Handling drift and failed deployments

    • Design and implement reusable Terraform modules for scalable infrastructure

    • Troubleshoot and resolve Terraform apply failures and infrastructure inconsistencies in production



  • Monitor system health using Prometheus, Grafana, and logging tools, and proactively identify issues

  • Perform root cause analysis (RCA) for production incidents and implement long-term fixes

  • Handle on-call incidents (24/7 rotation) and take full ownership until resolution

  • Work closely with development teams to improve system reliability, performance, and scalability

  • Automate operational tasks and improve deployment and infrastructure processes

  • Ensure security best practices across infrastructure, networking, and access controls

  • .

 




KIBO PERKS 






  • Flexible schedule and hybrid work setting 




  • Paid company holidays and global volunteer holiday 




  • Generous health, wellness, benefits, and time away programs 










  • Commitment to individual growth and development and opportunity for internal mobility 




  • Passionate, high-achieving teammates excited to help you succeed and learn 




  • Company-sponsored events and other activities  






At Kibo we celebrate and support all differences. Kibo is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital, disability, and veteran status. 




 



About Kibo

Composable digital commerce platform for eCommerce and order management.

Year founded
2016
Employees
350
Organization type
Private
Latest investment
Raised $21.60M Private Equity (2016) — led by Vista Equity Partners
Headquarters
US

Similar jobs

DevOps Engineer roles in Maharashtra
19h
Save
Mark Applied
Hide
DevOps Engineer
Pune, Maharashtra, India
HybridFull Time
Orion Innovation: Provides digital transformation, software engineering, and cloud services.
5+ YOERequires 5+ years of DevOps engineering experience, enterprise CI/CD implementation, cloud deployments, scripting, automation, application onboarding, DevSecOps familiarity, and strong troubleshooting and analytical skills.
Azure DevOps, GitHub Actions, Jenkins, GitLab CI/CD, Git, PowerShell, Bash, Python, Docker, Kubernetes, Terraform, Azure Resource Manager (ARM), CloudFormation, Ansible, Azure, AWS, REST APIs, YAML, Dynatrace, Splunk, AppDynamics, OpenTelemetry, ServiceNow, Jira, Confluence
1d
Save
Mark Applied
Hide
DEVOPS ENGINEER L3
Pune, Maharashtra, India
OnsiteFull Time
Wipro
WiproNYSE: WIT: Global technology services and consulting for digital transformation.
3+ YOERequires 3–5 years of DevOps experience, including continuous integration and deployment, pipeline management, infrastructure administration, tool configuration, automation, scripting, customer support, troubleshooting, and root cause analysis.
DevOps, continuous integration (CI), continuous deployment (CD)
1d
Save
Mark Applied
Hide
DevOps Engineer
Pune, Maharashtra, India
OnsiteFull Time
Accenture
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
2+ YOERequires 2+ years of Microsoft Azure DevOps experience and 15 years of full-time education. Requires cloud infrastructure, CI/CD, container orchestration, security, automation, troubleshooting, and collaboration skills.
Microsoft Azure DevOps, CI/CD
1d
Save
Mark Applied
Hide
DevOps Engineer (Azure) - 2- 6 years (Linux, scripting, AKS, Helm, Docker)
Pune, Maharashtra, India
HybridFull Time
NielsenIQ
NielsenIQNYSE: NIQ: Provides consumer intelligence and retail measurement data services.
2+ YOERequires 2–5 years in DevOps or system administration, Linux administration, scripting, Azure, AKS, Docker, Helm, CI/CD pipelines, troubleshooting, and 24x7 production operations; a BS in Computer Science or comparable program is required.
Linux, Bash, Python, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure ADO CI/CD pipelines, Docker, Helm, Terraform, ARM templates, Azure Monitor, Prometheus, Grafana, GitHub Copilot, GenAI
1d
Save
Mark Applied
Hide
Developer II - DevOps Engineering
Pune, Maharashtra, India
OnsiteFull Time
UST
UST: Global provider of digital transformation and IT services.
7+ YOEBachelor's degree or equivalent experience; 7–10 years in software, systems, databases, or networking; 4+ years with public cloud; GCP, AWS, Terraform, Jenkins, scripting, monitoring, and CI/CD experience.
AWS, Google Cloud Platform (GCP), Terraform, Jenkins, Google Cloud CLI, Google Cloud SDK, AppDynamics, CI/CD, Datadog, Google Cloud Monitoring, Shell, Python, PagerDuty, Git, Helm, GCE, GKE, IAM
1d
Save
Mark Applied
Hide
DevOps Engineer
Mumbai, Maharashtra, India
OnsiteFull Time
Sia
Sia: Provide management and AI consulting for business transformation.
4+ YOEEngineering background with 4+ years of DevOps experience; Python, Docker, container orchestration, cloud services, and CI/CD expertise required. English fluency required; Terraform and another programming language are advantageous.
Python, Docker, Kubernetes, Terraform, AWS, Google Cloud Platform (GCP), Microsoft Azure
1d
Save
Mark Applied
Hide
Senior Software Engineer - Devops
Pune, Maharashtra, India
OnsiteFull Time
DataDirect Networks
DataDirect Networks: High-performance storage and data management for AI and HPC.
7+ YOERequires a bachelor's or master's degree, 7+ years of Python software development in Linux, Kubernetes administration, containerized environments, Git, REST APIs, automation, and scripting.
Kubernetes, Argo CD, Harbor, Grafana, Alertmanager, Python, Linux, Git, REST APIs
2d
Save
Mark Applied
Hide
Associate DevOps Engineer (Observability)
Airoli Navi Mumbai, Maharashtra, India
RemoteFull Time
Teradata
TeradataNew York Stock Exchange: TDC: Provides cloud data analytics and enterprise AI software platforms.
1+ YOERequires 1+ year Grafana or observability administration experience, 1–3 years in DevOps or SRE, cloud provider experience, and experience with Terraform, Python, Ansible or Puppet, Jenkins or Bamboo, Git, Jira, SQL/noSQL, and Linux.
AWS, Azure, Google Cloud, Grafana, ServiceNow, Terraform, Python, Ansible, Puppet, Jenkins, Bamboo, Git, Jira, SQL, Linux