StarTree
Posted 1w ago

Software Engineer SRE

StarTree
India
RemoteFull Time
Responsibilities
  • monitoring services
  • troubleshooting issues
  • supporting deployments
Requirements
  • Requires 1–2 years in SRE
  • DevOps
  • Cloud operations
  • Systems engineering, or infrastructure support
  • Kubernetes
  • Cloud platform, Linux
  • Docker
  • Scripting
  • Monitoring, and CI/CD experience
Technical tools mentioned
KubernetesConfigMapsDockerLinuxBashPythonAWSGCPAzureCI/CDDNSHTTP/HTTPSTCP/IPApache Pinot

Job description

At StarTree we're a group of passionate individuals that desire to improve the lives of many by developing tools and technologies that support availability and speed in the world of real-time analytics. 

Our aim is to make it simple for every company to delight their users - external and internal - and create new revenue streams from their data, by building the world’s most comprehensive and accessible cloud analytics system.

We are looking for an SRE / Cloud Operations Engineer with 1–2 years of hands-on experience supporting cloud-based applications. The person will help maintain reliable, secure, and scalable platforms across AWS, GCP, and Azure, with a strong focus on Kubernetes-based workloads.

The role involves monitoring services, troubleshooting production issues, supporting deployments, automating operational tasks, and collaborating with development and platform teams to improve availability and performance.

Team Size

The candidate will join the SRE team, partnering closely with application developers, security teams, and cloud infrastructure teams. The team owns platform availability, observability, incident response, deployment reliability, and operational automation.

Candidate Experience

  • 1–2 years of experience in SRE, DevOps, Cloud Operations, Systems Engineering, or Infrastructure Support.
  • Hands-on experience working with Kubernetes, including deploying and troubleshooting pods, deployments, services, namespaces, ConfigMaps, and Secrets.
  • Exposure to supporting production applications and handling incidents or service requests.
  • Experience working in at least one public cloud: AWS, GCP, or Azure.
  • Hands-on exposure to Linux administration, scripting, monitoring, containers, and Docker.
  • Basic understanding of CI/CD pipelines and software deployment practices.

Skills (Non-negotiables & good to have)

  • Hands-on experience with at least one cloud platform: AWS, GCP, or Azure.
  • Basic Kubernetes knowledge: pods, deployments, services, namespaces, ConfigMaps, Secrets, logs, and troubleshooting.
  • Docker/container fundamentals.
  • Linux troubleshooting and administration skills.
  • Scripting knowledge in Bash or Python.
  • Understanding of networking basics: DNS, HTTP/HTTPS, TCP/IP, load balancers, security groups/firewalls.
  • Familiarity with monitoring and alerting concepts, such as metrics, logs, dashboards, and incident response.
  • Good communication and problem-solving skills.

About StarTree:  

StarTree is a cloud-based software company that enables business customers to derive advanced insights from real-time and historical data. StarTree was founded by the core software engineering team and inventors of Apache Pinot, which currently powers hundreds of user-facing applications at companies across industries, including LinkedIn, Uber, Target, 7Eleven, Etsy, Walmart, WePay, Factual, Weibo, and more. StarTree Cloud has enabled even more companies to deploy and operate real-time analytics at scale, including Stripe, Sovrn, Roadie, Just Eat Takeaway.com, Dialpad, Guitar Center, Blinkit, and more.


StarTree recently announced our Series B Funding with investment from GGV Capital, Sapphire Ventures, Bain Capital Ventures, and CRV. We have been named one of The Information's 50 Most Promising Startups and one of CRN's 10 Coolest Cloud Computing Startup Companies of 2022!

About StarTree

Private SaaS providing real-time analytics software for organizations using Apache Pinot.

Similar jobs

Site Reliability Engineer roles
7h
Save
Mark Applied
Hide
Site Reliability Engineer - Vice President
Pune, Maharashtra, India
HybridFull Time
Citi
CitiNYSE: C: Global financial services organization enabling growth and economic progress.
13+ YOE13+ years in SRE, production management, or software development; expertise in resilience, disaster recovery, distributed systems, OpenShift/Kubernetes, observability, IaC, automation, and Helm.
OpenShift, Kubernetes, Prometheus, Grafana, Loki, Mimir, Tempo, AppDynamics, Ansible, Terraform, Helm, Google Cloud, AWS, Azure, Java, Python, Go
10h
Save
Mark Applied
Hide
Site Reliability Engineer
India or Bengaluru
RemoteFull Time
Intelex
IntelexNYSE: FTV: Industrial technology providing essential, mission-critical workflow solutions.
5+ YOERequires 5+ years in SRE or system administration, 3+ years with CI/CD and cloud technologies, 1+ year with Docker/Kubernetes, AWS expertise, and a bachelor's degree or college diploma in a related field.
Windows, Linux, Terraform, ARM Templates, Cloud Formation, GitHub Actions, Octopus, Ansible, Jenkins, Azure DevOps, AWS, New Relic, Application Insights, AppDynamics, DataDog, Kubernetes, Bash, PowerShell, Python, Azure Automation Run book, SQL Server, IIS, CloudTrail, Docker, Scrum, Kanban, Lean, VNET, VNET peering, OAuth, AzureAD, ASE, ASP, AKS, Azure Apps, Load Balancers, Application Gateway, Firewall, API Management, SAP, PeopleSoft
10h
Save
Mark Applied
Hide
Site Reliability Engineer
India or Bengaluru or Mumbai
RemoteFull Time
Fortive Corporation
Fortive CorporationNYSE: FTV: Industrial technology providing essential, mission-critical workflow solutions.
5+ YOERequires 5+ years in SRE or system administration, 3+ years with CI/CD and cloud technologies, 1+ year with Docker/Kubernetes, AWS expertise, and a bachelor's degree or college diploma in a related field.
Windows, Linux, Terraform, ARM Templates, CloudFormation, GitHub Actions, Octopus, Ansible, Jenkins, Azure DevOps, AWS, IIS, CloudTrail, New Relic, Application Insights, AppDynamics, DataDog, Kubernetes, Bash, PowerShell, Python, Azure Automation Runbook, SQL Server, VNET, private link, Docker, OAuth, Azure AD, ASE, ASP, AKS, Azure Apps, Load Balancers, Application Gateway, Firewall, API Management, Scrum, Kanban, Lean, ITIL
12h
Save
Mark Applied
Hide
Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)
Bangalore, Karnataka, India
HybridFull Time
Zscaler
ZscalerNASDAQ: ZS: Cloud-native Zero Trust cybersecurity platform for digital transformation.
4+ YOERequires 4+ years designing, analyzing, and troubleshooting distributed systems, with hands-on SRE, Python, Terraform, Ansible, networking, Kubernetes, AWS, observability, web security, and DevOps experience.
Python, Terraform, Ansible, Kubernetes, AWS, HTTP, SSL/TLS, DNS, SQL, Grafana, CI/CD, Source Control Management (SCM), EKS, GKE, Linux, BSD
19h
Save
Mark Applied
Hide
Site Reliability Engineer - Vice President
Pune, Maharashtra, India
HybridFull Time
Citi
CitiNYSE: C: Global financial services organization enabling growth and economic progress.
13+ YOERequires 13+ years in SRE, production management, or software development; expertise in resilience, distributed systems, OpenShift/Kubernetes, observability, IaC, automation, and Helm.
OpenShift, Kubernetes, Prometheus, Grafana, Loki, Mimir, Tempo, AppDynamics, Ansible, Terraform, Helm, Google Cloud, AWS, Azure, Java, Python, Go
19h
Save
Mark Applied
Hide
Site Reliability Engineer - Vice President
Pune, Maharashtra, India
HybridFull Time
Citi
CitiNYSE: C: Global financial services organization enabling growth and economic progress.
13+ YOERequires 13+ years in SRE or production management, expertise in resilience, disaster recovery, distributed systems, Kubernetes/OpenShift, observability, IaC, automation, and Helm; strong strategic communication skills.
OpenShift, Kubernetes, Prometheus, Grafana, Loki, Mimir, Tempo, AppDynamics, Ansible, Terraform, Helm, Google Cloud, AWS, Azure, Java, Python, Go
2d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Hyderabad or Boca Raton or Asia
HybridFull Time
ModMed
ModMed: Private U.S. healthcare technology providing specialty-specific EHR, practice management, billing, and AI software to medical practices.
10+ YOERequires 10–13+ years in SRE or cloud architecture, advanced AWS, Kubernetes, DataDog, Terraform or Ansible, Python or Bash scripting, and executive-level technical communication.
AWS, Amazon EC2, AWS Lambda, Amazon RDS, Amazon S3, DataDog, Jenkins, Kubernetes, Terraform, Ansible, Python, Bash, Kafka, Amazon Kinesis, Amazon Redshift
2d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Hyderabad or Boca Raton
HybridFull Time
ModMed
ModMed: Private U.S. healthcare technology providing specialty-specific EHR, practice management, billing, and AI software to medical practices.
10+ YOE10–13+ years in SRE or cloud architecture; AWS, Kubernetes, DataDog, Terraform/Ansible, Python/Bash, scalable cloud-native systems, executive communication, and architectural leadership.
AWS, Amazon EC2, AWS Lambda, Amazon RDS, Amazon S3, Kubernetes, DataDog, Terraform, Ansible, Python, Bash, Jenkins, Kafka, Amazon Kinesis, Amazon Redshift