StoneX
Posted 1mo ago

Lead Engineer - Reliability Engineering

StoneX
Bengaluru, Karnataka, India
HybridFull Time
Responsibilities
  • defining standards
  • improving reliability
  • mentoring engineers
Requirements
  • 7+ years SRE/platform engineering experience
  • Proven track record implementing SLOs/SLIs
  • Observability
  • Incident management
  • Automation
  • Cloud and platform tooling
  • Mentorship and cross-team influence
Technical tools mentioned
DatadogTerraformKubernetesLinuxGitCI/CD pipelines

Job description

Overview:

Connecting clients to markets – and talent to opportunity.

With 5,400+ employees and over 80,000 institutional, commercial, and payments clients, we operate from more than 80 offices spread across six continents. As a Fortune 100, Nasdaq-listed provider, we connect clients to the global markets – focusing on innovation, human connection, and providing world-class products and services to all types of investors.

Whether you want to forge a career connecting our retail clients to potential trading opportunities, or ingrain yourself in the world of institutional investing, StoneX Group is made up of four business segments that offer endless potential for progression and growth.

 

Engage in a deep variety of business-critical activities that keep our company running efficiently. From strategic marketing and financial management to human resources and operational oversight, you’ll have the opportunity to optimize processes and implement game-changing policies.

 

As a Lead Engineer in Reliability Engineering, you will help define and drive the next stage of reliability maturity across our platforms and services. This is a senior hands-on engineering role for someone who has already spent several years building and operating Site Reliability Engineering practices in a large organization, and who understands what good looks like in production at scale.

You will partner closely with our Platform Engineering Observability team to improve reliability standards, operational practices, service ownership models, and engineering guardrails. Over time, you will help grow reliability capabilities across the wider engineering organization by mentoring engineers, shaping ways of working, and building practical and measurable reliability practices. The team is actively expanding end-to-end observability coverage across key business applications. This role will help ensure that telemetry, service health, reliability standards, and operational practices are implemented consistently and effectively as adoption grows.

This is an individual contributor role with no direct people management responsibilities. Success in this role will come through technical leadership, hands on engineering contribution, mentorship, and influence across teams.

 



Responsibilities:

 

  • Define and drive reliability engineering standards, practices, and the enterprise reliability maturity model across platforms and services, including service tiering and adoption metrics
  • Partner with engineering, platform, infrastructure, and product teams to improve service reliability, resilience, operability, and supportability
  • Establish and mature reliability practices such as SLOs, SLIs, error budgets, alert quality, toil reduction, production readiness reviews, and service ownership expectations
  • Build and improve operational processes for change safety, release confidence, capacity planning, resilience testing, and disaster recovery across multi-cloud and hybrid environments
  • Use observability platforms such as Datadog or similar tools to improve visibility, actionable alerting, dashboards, and service health reporting
  • Drive end-to-end observability adoption for critical applications, ensuring consistent implementation of metrics, logs, traces, service maps, dashboards, and actionable alerting
  • Partner with application, platform, and infrastructure teams to improve instrumentation quality, service ownership, and operational readiness as key applications are onboarded into the observability ecosystem
  • Define and standardize observability architecture and telemetry standards, including service dependency mapping, service health indicators, alert quality, and operational response workflows
  • Drive automation of operational tasks and embed reliability guardrails into platform engineering workflows, including CI/CD pipelines and internal developer platforms
  • Apply observability standards to golden path templates, workflows, scorecards, and dashboards within the internal developer platform so operational best practices are embedded holistically throughout the SDLC
  • Identify reliability risks including architectural weaknesses, service fragility, and third-party or provider dependencies, and partner with teams to address them
  • Define and track meaningful reliability metrics and operational KPIs that help engineering teams improve service outcomes over time
  • Act as a senior hands-on engineer who guides technical direction while contributing directly to design, implementation, and operational improvement work
  • Mentor and coach engineers across the team, helping them develop stronger reliability and operational engineering skills


Qualifications:

 

  • A track record of building, improving, or scaling reliability engineering or SRE practices in a large organization
  • 7+ years of experience in SRE, production engineering, platform engineering, infrastructure engineering, or a closely related role
  • Several years of hands-on experience supporting production systems at scale, including incident response, problem management, availability improvement, and operational excellence
  • Strong practical experience defining and implementing SLOs, SLIs, error budgets, service health models, and reliability focused engineering practices
  • Strong experience with observability platforms such as Datadog or similar platforms, including metrics, logs, tracing, alerting, dashboards, and service level reporting
  • Experience driving or supporting end-to-end observability adoption across application teams, including instrumentation, telemetry standards, dashboards, alerting, and service level reporting
  • Experience driving improvements in incident management, post incident reviews, on call effectiveness, and operational maturity
  • Experience automating operational processes using tools such as Terraform, scripting languages, CI and CD pipelines, and cloud native platforms
  • Experience working with Kubernetes, Linux, Git, and modern cloud or platform infrastructure
  • Strong systems thinking and the ability to balance reliability, latency, engineering velocity, risk, and cost
  • Strong communication, collaboration, and influencing skills, with the ability to work across multiple teams and levels of seniority
  • Demonstrated ability to mentor engineers and help raise the reliability maturity of a broader team
  • A practical mindset, someone who can define strong engineering practices and also contribute directly in a hands-on way

 

What makes you stand out:

 

  • You have helped build or formalize reliability engineering or SRE practices in a complex organization
  • You know what good looks like for service ownership, production readiness, alerting quality, incident response, and operational accountability
  • You have helped onboard critical applications into an end-to-end observability model, improving visibility across metrics, logs, traces, service dependencies, and operational response
  • You have successfully reduced toil, improved service reliability, and created measurable operational improvements across teams
  • You are able to influence engineering culture, not just tooling or process
  • You have coached less experienced engineers and helped teams grow into stronger operational ownership
  • You are comfortable introducing structure and standards without creating unnecessary bureaucracy
  • You can work across observability, platform engineering, and application teams to create practical, adoptable reliability practices

 

Education / Certification Requirements: 

 

  • Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience
  • Relevant certifications are a plus, but practical experience building and operating reliable systems at scale is more important
  • Commitment to continual professional and technical development

 

Working environment:

  • Hybrid, four days in the office.
  • Occasional Travel Requirements, for team collaboration meetings and conferences.

 

 

#LI-Hybrid

About StoneX

Provides financial services, global market access, and trading platforms.

Similar jobs

Reliability Engineer roles near Bengaluru, Karnataka
1w
Save
Mark Applied
Hide
Reliability Engineer
Bengaluru or Gurugram or North America
OnsiteFull Time
JLL
JLLNYSE: JLL: Global commercial real estate and investment management services.
5+ YOEBachelor's degree in engineering, preferably mechanical or electrical, and 5–10 years implementing RCM, condition-based and predictive maintenance, building automation, asset management, analytics, and capital planning.
Building Automation Systems (BAS), Computerized Maintenance Management System (CMMS), Microsoft Excel, automated fault detection and diagnostics, Reliability Centered Maintenance (RCM), condition-based maintenance (CbM), predictive maintenance, predictive testing and inspection (PT&I), Failure Modes and Effects Analysis (FMEA)
1w
Save
Mark Applied
Hide
Reliability Engineer
Bangalore, Karnataka, India
HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
8+ YOEEngineering bachelor's degree with 8–10 years' experience or master's degree with 6–8 years' experience; expertise in hardware reliability, data center infrastructure, network architecture, PCBA failures, analytics, and project leadership.
FMEA, RDT, Reliability, Availability, and Serviceability (RAS), ORT, CLCA
1w
Save
Mark Applied
Hide
Reliability Engineer - (REE)
Bengaluru or Hyderabad
OnsiteFull Time
London Stock Exchange Group
London Stock Exchange GroupLondon Stock Exchange: LSEG: Provides financial market infrastructure and global data analytics services.
Enterprise application support experience, ITSM knowledge, ticketing tools, incident diagnosis, log analysis, SLA adherence, strong communication, and problem-solving skills.
AWS, ServiceNow, Jira Service Management, Remedy, APIs, SSO, Azure AD, MFA, Microsoft 365, Clarity PPM, Condeco, Cornerstone, SuccessFactors Learning
3w
Save
Mark Applied
Hide
Reliability Engineer
Bengaluru, Karnataka, India
OnsiteFull Time
Quest Global
Quest Global: Global engineering services for product development and lifecycle management.
10+ YOEBE/B.Tech in Mechanical and 10+ years experience; lead technical writers; develop SOPs, operating and maintenance manuals; convert PFDs/P&IDs into user-focused content; coordinate multidisciplinary teams for refinery and petrochemical projects.
CMMS
4w
Save
Mark Applied
Hide
Delhi Career Fair 5 & 6 Sept - Reliability Engineer
Bengaluru, Karnataka, India
OnsiteFull Time
ExxonMobil
ExxonMobilNYSE: XOM: Produces and distributes oil, natural gas, and petrochemical products.
3+ YOEBachelor's in engineering, minimum 3 years reliability or engineering experience in manufacturing/oil & gas, knowledge of reliability concepts, statistical methods, strong communication and problem-solving skills.
1mo
Save
Mark Applied
Hide
Staff Reliability Engineer (System-Level Modeling — SOFC Systems)
Bengaluru, Karnataka, India
OnsiteFull Time
Bloom Energy
Bloom EnergyNYSE: BE: Manufactures solid oxide fuel cell systems for onsite power.
4+ YOEBachelor's degree in engineering, 4+ years reliability experience with system-level RBD/Markov/Monte Carlo modeling, proficiency with ReliaSoft BlockSim, strong statistical and communication skills.
ReliaSoft BlockSim, Python, MATLAB
1mo
Save
Mark Applied
Hide
Lead Engineer - Reliability Engineering
Bangalore, Karnataka, India
HybridFull Time
StoneX Group
StoneX GroupNASDAQ: SNEX: Provides global financial services and market access.
7+ YOE7+ years SRE/platform experience, hands-on production reliability, SLO/SLI/error budget expertise, observability and Datadog experience, automation with Terraform/CI/CD, Kubernetes/Linux/Git knowledge, and ability to mentor engineers.
Datadog, Terraform, Kubernetes, Linux, Git, CI/CD
1mo
Save
Mark Applied
Hide
Reliability Engineer (Bangalore, IN, 560071)
Bengaluru, Karnataka, India
OnsiteFull Time
SBM Offshore
SBM OffshoreEuronext Amsterdam: SBMO: Provides floating production solutions for the offshore energy industry.
2+ YOEBachelor's in engineering, minimum 2 years maintenance management experience, strong CMMS and maintenance planning knowledge, incident investigation and RCA experience, KPI development and audit implementation; oil & gas experience desired.
Computerized Maintenance Management System (CMMS)