Blackstone
Posted 6mo ago

Site Reliability Engineer - Data, Cloud & Developer Experience

Blackstone
New York, New York, United States
$140k-$225k/yrOnsiteFull Time
Responsibilities
  • Lead SRE
  • Observability integration
  • Standard evolution
Requirements
  • 5+ years in infrastructure/software/devops/platform engineering
  • Strong AWS
  • Terraform
  • Docker
  • Kubernetes
  • Observability tools
  • Coding in Python, C#, TypeScript
  • Automation
  • Incident management
Technical tools mentioned
AWSTerraformPuppetGitLab CIDockerKubernetesGrafanaPrometheusSplunkPythonC#TypeScript

Job description

Blackstone is the world’s largest alternative asset manager. We seek to create positive economic impact and long-term value for our investors, the companies we invest in, and the communities in which we work. We do this by using extraordinary people and flexible capital to help companies solve problems. Our $1.1 trillion in assets under management include investment vehicles focused on private equity, real estate, public debt and equity, infrastructure, life sciences, growth equity, opportunistic, non-investment grade credit, real assets and secondary funds, all on a global basis. Further information is available at www.blackstone.com. Follow @blackstone on LinkedInX, and Instagram.

Role:    

Blackstone's Site Reliability Engineering team is responsible for improving the reliability of systems and services to meet the needs of the business. This is achieved through collaboration with the development and engineering teams to leverage SRE practices and principles. You'll have the opportunity to identify and solve new problems as they arise, deploy and maintain observability systems and pipelines, enhance operations and support for services and platforms, and pursue emerging opportunities for efficiency and business value. This position involves the selection, implementation, and maintenance of key observability tooling. It requires ongoing evaluation of the firms needs in observability, monitoring, alerting, resilience, and recovery.

We work alongside service owners on design, implementation, and management of services for continuous improvement. We achieve the requisite reliability of services using clear definitions and measurable targets. We plan for and practice recovery from disaster scenarios and respond in real time to incidents. We guide the postmortem process in order to mitigate risks, prevent future disruptions, and improve the on-call experience. We aim to eliminate manual work, improve operational efficiency, and ensure high-quality outputs in all that we do.

Key Responsibilities:

  • Provide technical leadership in the understanding and adoption of SRE methodologies across the firm      

  • Incorporate observability standards into code and deployment pipelines

  • Evolve the SRE standards that are adopted across all teams

  • Partner with colleagues in various roles and reporting lines to improve service reliability and operational efficiency

  • Assist developers and engineers directly and through AI assistants

  • Implement instrumentation and provide comprehensive performance insights to service owners

  • Ensure monitoring and alerting reflects the reliability of services for users and enables effective on-call operations

  • Implement strategic observability tools and work to control overhead in maintenance and cost

  • Participate in on-call rotations and respond to system incidents to ensure service availability and minimize operational impact

  • Use automation to manage, maintain, and scale SRE systems with minimal human intervention

  • Foster a blameless team culture while assisting in postmortem discussions and reporting

Qualifications:

  • 5 + years of professional experience with either, Infrastructure Engineering, Software Engineering, DevOps Engineering or Platform Engineering.

  • Automation script writing skills; effectively reads and troubleshoots code (Python, C#, Typescript, etc.)

  • Makes effective use of coding assistants and chat models (Anthropic, OpenAI)

  • Proficiency with public cloud providers (strong AWS experience required, preferred Azure experience)

  • Configuration-as-code, infrastructure management, and CI/CD tooling experience (Terraform, Puppet, Gitlab CI)

  • Hand-on experience with Docker and container schedulers including AWS ECS & EKS

  • Excellent troubleshooting skills for Linux and Windows, and networking experience with observability tools (Grafana, Prometheus, Splunk, etc.)

  • Comfortable under pressure with incident management and collaborating during postmortems

  • Excellent communication and organizational skills

  • Curiosity and motivation to improve systems and processes through a sense of shared ownership


The duties and responsibilities described here are not exhaustive and additional assignments, duties, or responsibilities may be required of this position.  Assignments, duties, and responsibilities may be changed at any time, with or without notice, by Blackstone in its sole discretion.

Expected annual base salary range:

$140,000 - $225,000

Actual base salary within that range will be determined by several components including but not limited to the individual's experience, skills, qualifications and job location. For roles located outside of the US, please disregard the posted salary bands as these roles will follow a separate compensation process based on local market comparables.

Additional compensation and benefits offered in connection with the role consist of comprehensive health benefits, including but not limited to medical, dental, vision, and FSA benefits; paid time off; life insurance; 401(k) plan; and discretionary bonuses. Certain employees may also be eligible for equity and other incentive compensation at Blackstone’s sole discretion.

Blackstone is committed to providing equal employment opportunities to all employees and applicants for employment without regard to race, color, creed, religion, sex, pregnancy, national origin, ancestry, citizenship status, age, marital or partnership status, sexual orientation, gender identity or expression, disability, genetic predisposition, veteran or military status, status as a victim of domestic violence, a sex offense or stalking, or any other class or status in accordance with applicable federal, state and local laws. This policy applies to all terms and conditions of employment, including but not limited to hiring, placement, promotion, termination, transfer, leave of absence, compensation, and training.  All Blackstone employees, including but not limited to recruiting personnel and hiring managers, are required to abide by this policy.

If you need a reasonable accommodation to complete your application, please contact Human Resources at 212-583-5000 (US), +44 (0)20 7451 4000 (EMEA) or +852 3656 8600 (APAC).

Depending on the position, you may be required to obtain certain securities licenses if you are in a client facing role and/or if you are engaged in the following:

  • Attending client meetings where you are discussing Blackstone products and/or and client questions;

  • Marketing Blackstone funds to new or existing clients;

  • Supervising or training securities licensed employees;

  • Structuring or creating Blackstone funds/products; and

  • Advising on marketing plans prepared by a sales team or developing and/or contributing information for marketing materials.

Note: The above list is not the exhaustive list of activities requiring securities licenses and there may be roles that require review on a case-by-case basis.  Please speak with your Blackstone Recruiting contact with any questions.

To submit your application please complete the form below. Fields marked with a red asterisk * must be completed to be considered for employment (although some can be answered "prefer not to say"). Failure to provide this information may compromise the follow-up of your application. When you have finished click Submit at the bottom of this form.

About Blackstone

Global alternative asset management firm.

Similar jobs

Site Reliability Engineer roles near New York, New York
1d
Save
Mark Applied
Hide
Site Reliability Engineer for CIAM
Whippany, New Jersey, United States
$120k-$175k/yr OnsiteFull Time
Barclays
BarclaysLondon Stock Exchange: BARC: Global bank providing retail, corporate, and investment financial services.
Experience designing and operating highly available cloud systems; AWS, Kubernetes, ECS, Python, Bash, JSON/YAML, disaster recovery, zero-downtime deployment, microservices, APIs, and SRE practices.
AWS, Azure, GCP, Kubernetes, ECS, Fargate, GCE, Python, Bash, JSON, YAML, ForgeRock, PingGateway, PingAM, PingIDM, PingDS, HTTP, PKI
1d
Save
Mark Applied
Hide
Site Reliability Engineer, SaaS
New York City, New York, United States
$145k-$175k/yr HybridFull Time
Columbus Blue Jackets
Columbus Blue Jackets: Professional ice hockey team competing in the National Hockey League.
5+ YOERequires 5+ years in systems reliability, SRE, or SaaS operations; expertise in Microsoft 365, Azure, Slack, Zoom, email flow, identity platforms, automation, scripting, and SLA/SLO/SLI monitoring.
Microsoft 365, Microsoft SharePoint, Microsoft Copilot, Azure, Slack, Adobe, Zoom, Entra ID, PowerShell, ServiceNow, SSO, SCIM, APIs, ITSM, SOC2
1d
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Playout
Stamford, Connecticut, United States
$145k-$175k/yr HybridFull Time
NBCUniversal
NBCUniversalNASDAQ: CMCSA: Produces and distributes entertainment content and theme park experiences.
8+ YOEBachelor's degree or equivalent experience, 8 years of engineering experience in broadcast playout, Linux administration, cloud and networking expertise, monitoring, containerization, and 24/7 on-call availability.
Linux, Splunk, Grafana, ServiceNow, Docker, Kubernetes, AWS, Snell, Harris, Imagine, Amagi, Evertz, GrassValley, Harmonic, CoralBay, Veset, TS, HEVC, H.264, HLS, CMAF, SCTE-35, SCTE-224, ESAM, SRT, RIST, Slack
2d
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Cloud, Observability & Automation)
Jersey City, New Jersey, United States
HybridFull Time
DTCC
DTCC: Provides post-trade infrastructure for the global financial services industry
8+ YOEBachelor's degree in computer science, engineering, or equivalent experience; 8+ years in SRE, production engineering, DevOps, or application support; AWS, Python, Java, Go, Linux, monitoring, incident management, and distributed systems expertise.
AWS, Splunk, Grafana, Dynatrace, ITSI, Python, Java, Amazon Q, Kiro, Go, Linux/Unix
2d
Save
Mark Applied
Hide
Staff+ Site Reliability Engineer, Safeguards ML Infra
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOEProduction change-management experience, high-stakes release and on-call experience, AWS/GCP operations, Python proficiency, and a bachelor's degree or equivalent experience.
Python, Rust, AWS, GCP, AWS Bedrock, GCP Vertex, Claude
2d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Paze
Scottsdale or Chicago or San Francisco or New York City or Phoenix or California or Illinois
$106k-$156k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
3+ YOEBachelor's degree in business, computer science, or related field; 3+ years of related technical or software development experience; Linux administration, Git, scripting, observability, incident management, and enterprise-scale experience required.
Linux, Git, Java, Ruby, Python, JavaScript, Go, AWS, Docker, Kubernetes, Swarm, CI/CD, TCP/UDP/IP
2d
Save
Mark Applied
Hide
SRE / Infrastructure Engineer (LABGEN)
Great Neck, New York, United States
$90k-$115k/yr OnsiteFull Time
Medfar
Medfar: Provides cloud-based electronic medical record software for healthcare clinics.
5+ YOERequires 5+ years in SRE, infrastructure, DevOps, or similar roles; Linux and Windows Server administration; Apache, networking, cloud, backups, disaster recovery, SQL, CI/CD, security, and English communication skills.
RHEL, CentOS, Ubuntu, Windows Server, Apache, Azure, AWS, GCP, SQL, .NET, JavaScript, Git, CVS, Jenkins, Prometheus, Grafana, ELK, Datadog, SentinelOne, Python, Bash, LIS, EHR, CI/CD
3d
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineer - Remote
Basking Ridge, New Jersey, United States
$113k-$193k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
10+ YOE5+ MgmtBachelor’s degree in a relevant field, 10+ years in software, SRE, platform, DevOps, infrastructure, or technology operations, and 5+ years leading engineering or operational teams. Requires production, cloud, ITSM, and reliability experience.
Azure, AWS, Infrastructure-as-Code, Kubernetes, OpenShift, Datadog, Splunk, Grafana, Prometheus, OpenTelemetry, AIOps, ChatOps, LLM