DTCC
Posted 5d ago

Principal Site Reliability Engineer (Cloud, Observability & Automation)

DTCC
Jersey City, New Jersey, United States
HybridFull Time
Responsibilities
  • driving reliability
  • designing observability
  • leading incident response
Requirements
  • Bachelor's degree in computer science
  • Engineering, or equivalent experience
  • 8+ years in SRE
  • Production engineering
  • DevOps, or application support
  • AWS
  • Python, Java, Go, Linux
  • Monitoring
  • Incident management, and distributed systems expertise
Technical tools mentioned
AWSSplunkGrafanaDynatraceITSIPythonJavaAmazon QKiroGoLinux/Unix

Job description

Are you ready to make an impact at DTCC?

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Pay and Benefits: 

  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits, based on location
  • Pension / Retirement benefits
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).

The Impact You Will Have in This Role

The Enterprise Application Support (EAS) team supports critical applications across the ITP and ECS business lines, ensuring the reliability, scalability, and performance of enterprise platforms.

As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across mission-critical systems. You will lead reliability initiatives, champion observability and automation, drive major incident response, and partner with engineering, infrastructure, and security teams to build resilient, highly available applications.

This is a hands-on technical leadership role focused on improving system performance, reducing operational risk, accelerating recovery, and advancing SRE best practices through modern cloud, observability, automation, and AI-powered technologies.

Your Primary Responsibilities:

  • Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications. 
  • Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms. 
  • Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs. 
  • Lead major incident response, root cause analysis, and continuous service improvement initiatives. 
  • Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies. 
  • Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle. 
  • Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives. 
  • Identify operational risks and deliver strategic reliability improvements across the technology ecosystem. 
  • Collaborate with technical and business stakeholders to improve service reliability and operational outcomes.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience. 
  • 8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines.

Talent Needed for Success

  • Strong hands-on experience with AWS and cloud-native architectures. 
  • Proficiency in Python, Java, Go, or similar programming languages. 
  • Strong Linux/Unix systems administration and troubleshooting experience. 
  • Expertise in observability and monitoring platforms including Splunk, Grafana, Dynatrace, and ITSI. 
  • Experience leading major incident management and root cause investigations in complex production environments. 
  • Strong understanding of distributed systems, resiliency engineering, performance tuning, automation, and operational excellence. 
  • Excellent communication and stakeholder management skills with the ability to influence technical and business partners.

Preferred Qualifications

  • Experience with AI-assisted engineering tools such as Amazon Q, Kiro, or similar technologies. 
  • Experience designing and measuring SLOs, SLIs, and operational KPIs. 
  • Experience supporting large-scale enterprise applications in financial services or other highly regulated environments.

The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

About Company

Serves as a dedicated technology resource for advancing DTCC’s business opportunities and providing industry thought leadership for leveraging new technology. The goal of this new department is to partner internally with IT, our business and regulatory divisions and externally with clients, regulators, and fintech vendors, to help build new platforms and business models to advance DTCC’s mission to support the financial markets.

Company


With over 50 years of experience, DTCC is the premier post-trade market infrastructure for the global financial services industry. From 20 locations around the world, DTCC, through its subsidiaries, automates, centralizes, and standardizes the processing of financial transactions, mitigating risk, increasing transparency, enhancing performance and driving efficiency for thousands of broker/dealers, custodian banks and asset managers. Industry owned and governed, the firm innovates purposefully, simplifying the complexities of clearing, settlement, asset servicing, transaction processing, trade reporting and data services across asset classes, bringing enhanced resilience and soundness to existing financial markets while advancing the digital asset ecosystem. In 2024, DTCC’s subsidiaries processed securities transactions valued at U.S. $3.7 quadrillion and its depository subsidiary provided custody and asset servicing for securities issues from over 150 countries and territories valued at U.S. $99 trillion. DTCC’s Global Trade Repository service, through locally registered, licensed, or approved trade repositories, processes more than 25 billion messages annually. To learn more, please visit us at www.dtcc.com or connect with us on LinkedIn, X, YouTube, Facebook and Instagram.

DTCC proudly supports Flexible Work Arrangements favoring openness and gives people freedom to do their jobs well, by encouraging diverse opinions and emphasizing teamwork.  When you join our team, you’ll have an opportunity to make meaningful contributions at a company that is recognized as a thought leader in both the financial services and technology industries. A DTCC career is more than a good way to earn a living. It’s the chance to make a difference at a company that’s truly one of a kind.

Learn more about Clearance and Settlement by clicking here.

About DTCC

Provides post-trade infrastructure for the global financial services industry

Similar jobs

Site Reliability Engineer roles near Jersey City, New Jersey
3d
Save
Mark Applied
Hide
Site Reliability Engineer, Global Banking & Markets, Vice President
New York City, New York, United States
$150k-$250k/yr OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
8+ YOE8+ years of software or reliability engineering experience; proficiency in a major programming language; cloud, distributed systems, SRE, automation, observability, incident response, and risk management expertise.
Java, GCP, AWS, Kubernetes, Docker, Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist, Apache Kafka, Spring Boot, gRPC, Protocol Buffers, Apache Camel, Spring Integration, Terraform, Helm, Prometheus, Grafana, OpenTelemetry, SQL, NoSQL, Vert.x, Netty, Microsoft?
5d
Save
Mark Applied
Hide
Site Reliability Engineer, SaaS
New York City, New York, United States
$145k-$175k/yr HybridFull Time
Columbus Blue Jackets
Columbus Blue Jackets: Professional ice hockey team competing in the National Hockey League.
5+ YOERequires 5+ years in systems reliability, SRE, or SaaS operations; expertise in Microsoft 365, Azure, Slack, Zoom, email flow, identity platforms, automation, scripting, and SLA/SLO/SLI monitoring.
Microsoft 365, Microsoft SharePoint, Microsoft Copilot, Azure, Slack, Adobe, Zoom, Entra ID, PowerShell, ServiceNow, SSO, SCIM, APIs, ITSM, SOC2
5d
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Playout
Stamford, Connecticut, United States
$145k-$175k/yr HybridFull Time
NBCUniversal
NBCUniversalNASDAQ: CMCSA: Produces and distributes entertainment content and theme park experiences.
8+ YOEBachelor's degree or equivalent experience, 8 years of engineering experience in broadcast playout, Linux administration, cloud and networking expertise, monitoring, containerization, and 24/7 on-call availability.
Linux, Splunk, Grafana, ServiceNow, Docker, Kubernetes, AWS, Snell, Harris, Imagine, Amagi, Evertz, GrassValley, Harmonic, CoralBay, Veset, TS, HEVC, H.264, HLS, CMAF, SCTE-35, SCTE-224, ESAM, SRT, RIST, Slack
6d
Save
Mark Applied
Hide
Staff+ Site Reliability Engineer, Safeguards ML Infra
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOEProduction change-management experience, high-stakes release and on-call experience, AWS/GCP operations, Python proficiency, and a bachelor's degree or equivalent experience.
Python, Rust, AWS, GCP, AWS Bedrock, GCP Vertex, Claude
6d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Paze
Scottsdale or Chicago or San Francisco or New York City or Phoenix or California or Illinois
$106k-$156k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
3+ YOEBachelor's degree in business, computer science, or related field; 3+ years of related technical or software development experience; Linux administration, Git, scripting, observability, incident management, and enterprise-scale experience required.
Linux, Git, Java, Ruby, Python, JavaScript, Go, AWS, Docker, Kubernetes, Swarm, CI/CD, TCP/UDP/IP
6d
Save
Mark Applied
Hide
SRE / Infrastructure Engineer (LABGEN)
Great Neck, New York, United States
$90k-$115k/yr OnsiteFull Time
Medfar
Medfar: Provides cloud-based electronic medical record software for healthcare clinics.
5+ YOERequires 5+ years in SRE, infrastructure, DevOps, or similar roles; Linux and Windows Server administration; Apache, networking, cloud, backups, disaster recovery, SQL, CI/CD, security, and English communication skills.
RHEL, CentOS, Ubuntu, Windows Server, Apache, Azure, AWS, GCP, SQL, .NET, JavaScript, Git, CVS, Jenkins, Prometheus, Grafana, ELK, Datadog, SentinelOne, Python, Bash, LIS, EHR, CI/CD
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineer - Remote
Basking Ridge, New Jersey, United States
$113k-$193k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
10+ YOE5+ MgmtBachelor’s degree in a relevant field, 10+ years in software, SRE, platform, DevOps, infrastructure, or technology operations, and 5+ years leading engineering or operational teams. Requires production, cloud, ITSM, and reliability experience.
Azure, AWS, Infrastructure-as-Code, Kubernetes, OpenShift, Datadog, Splunk, Grafana, Prometheus, OpenTelemetry, AIOps, ChatOps, LLM
1w
Save
Mark Applied
Hide
Site Reliability Engineer
New York or Hong Kong or London or Singapore
$105k-$300k/yr OnsiteFull Time
Citadel
Citadel: Global alternative investment management firm
Bachelor's degree in computer science, related STEM discipline, or equivalent experience; proficiency in a modern structured programming language, software development practices, distributed systems, and strong communication skills.
Python, SQL, JavaScript, CSS, React, CI/CD