Nscale
Posted 7h ago

Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale
Houston or New York City or San Francisco or Seattle
OnsiteFull Time
Responsibilities
  • leading programs
  • tracking metrics
  • optimizing workflows
Requirements
  • 5+ years in technical program management for complex infrastructure or software programs
  • Knowledge of data centers
  • Distributed systems, Linux
  • Networking
  • Operational metrics, and Agile/Scrum. Technical degree preferred
Technical tools mentioned
LinuxAgileScrumNVIDIA GPUsInfiniBandRDMASREContinuous Integration/Continuous Deployment (CI/CD)

Job description

Role Overview

As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will be responsible for driving complex, cross-functional programs that ensure the stability, availability, and growth of our cutting-edge GPU fleet and Infiniband network fabrics. This role requires a blend of deep technical understanding, rigorous program management, and a relentless focus on delivering against key operational metrics (SLAs, Uptime, Availability). You will bridge the gap between engineering execution and strategic business goals, directly impacting our ability to serve customer workloads at scale.

Key Responsibilities

  • Program Leadership: Own the planning, execution, and delivery of strategic operational programs, including new data center AI infrastructure build-outs, large-scale fleet software/firmware rollouts, and the implementation of new operational tooling (in partnership with SRE).
  • Metrics and Reporting: Establish, track, and drive accountability against critical infrastructure KPIs, specifically focusing on Availability (Target 97.5%) and Uptime (Target 99%). Develop clear dashboards and communication rhythms to provide leadership with real-time visibility into operational health, program status, and risk.
  • Process Engineering: Analyze and optimize operational workflows across Fleet Operations, Network Operations, and SRE teams. Drive the standardization of incident management, change management, and postmortem processes to reduce toil and improve Mean Time to Recovery (MTTR).
  • Cross-Functional Coordination: Serve as the primary liaison between engineering teams (Hardware, Compute Platform, Network), Data Center Operations, and external vendors (GPU, Network hardware). Proactively identify and resolve dependencies, risks, and roadblocks.
  • Capacity and Readiness: Partner with Data Science/Operation Programs to translate capacity planning models into actionable infrastructure delivery and readiness roadmaps. Ensure that new hardware (GPUs, NICs, switches) is successfully integrated into the operational control plane and meets go-live criteria.
  • Risk Management: Proactively identify technical, schedule, and resource risks related to AI infrastructure scaling and stability. Develop mitigation strategies and communicate impacts clearly to stakeholders.

Required Qualifications

  • Experience: 5+ years of experience in a Technical Program Management role, successfully driving large-scale, complex infrastructure or software engineering programs.
  • Technical Domain Knowledge: Strong foundational understanding of data center infrastructure, distributed systems, Linux, and networking concepts.
  • Program Management Rigor: Proven expertise in modern program management methodologies (Agile, Scrum, PMP certification preferred). Exceptional organizational, communication, and presentation skills.
  • Metrics-Driven Approach: Demonstrable experience in defining, tracking, and improving system performance based on operational metrics (e.g., Uptime, Availability, MTTR, SLOs/SLIs).
  • Execution in Ambiguity: Ability to thrive in a fast-paced, high-growth environment, managing multiple priorities and adapting to evolving technical requirements.

Preferred Qualifications

  • Direct experience managing programs related to data center infrastructure build-outs and hardware commissioning processes.
  • Specific domain knowledge of AI/HPC infrastructure, including NVIDIA GPUs, InfiniBand/RDMA networks, and the challenges of tightly-coupled systems.
  • Experience in a hyperscale or public cloud environment supporting 24/7 mission-critical services.
  • Familiarity with SRE principles, automation tooling, and continuous integration/continuous deployment (CI/CD) pipelines for infrastructure.
  • A Bachelor's or Master's degree in a technical field (Computer Science, Engineering, etc.) or equivalent practical experience.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package (base + equity) with reviews every 12 months. 
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI. 
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

About Nscale

Vertically integrated AI infrastructure provider for high-performance computing.

Year founded
2023
Employees
400
Organization type
Private
Latest investment
Raised $790.00M Debt Financing (2026)
Headquarters
GB

Similar jobs

Technical Program Manager roles near Houston, Texas
4d
Save
Mark Applied
Hide
Technical Program Manager
Spring or Houston
$93k-$144k/yr OnsiteFull Time
HP
HPNYSE: HPQ: Manufactures personal computers, printers, and 3D printing hardware.
5+ YOEBachelor's or master's degree in business management, engineering, or computer science; 5–7 years of experience; program leadership, technical knowledge, stakeholder management, communication, and analytical skills.
AMD, Intel, Qualcomm, Microsoft, Google
4d
Save
Mark Applied
Hide
Technical Program Manager
Spring, Texas, United States
$93k-$144k/yr OnsiteFull Time
HP
HPNYSE: HPQ: Manufacturer of personal computers, printers, and imaging devices.
5+ YOEBachelor's or master's degree in business management, engineering, or computer science; typically 5–7 years of experience; program planning, technical leadership, stakeholder management, communication, analytical, and problem-solving skills.
AMD, Intel, Qualcomm, Microsoft, Google
6d
Save
Mark Applied
Hide
Lead Technical Program Manager
Houston, Texas, United States
OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOERequires 5+ years of technical program management or equivalent expertise, complex technology delivery experience, AI-assisted workflow validation, stakeholder management, resource and budget management, and advanced analytical reasoning.
AI
6d
Save
Mark Applied
Hide
Technical Program Manager - Houston, TX
Houston, Texas, United States
$110k-$130k/yr FieldFull Time
SET Environmental
SET Environmental: Provides nationwide hazardous waste management and environmental remediation services.
5+ YOEBachelor's degree or equivalent experience; 5-8 years in environmental or hazardous waste services; knowledge of RCRA, DOT, OSHA, HAZCOM, waste profiling, and manifesting; regional travel required.
Microsoft Office
6d
Save
Mark Applied
Hide
Technical Program Manager — Data Center Delivery USA
Houston or North America or United States
HybridFull Time
Submer
Submer: Sustainable liquid immersion cooling systems for data centers.
7+ YOEBachelor's degree or equivalent experience; 7+ years in technical programs, projects, engineering, construction, commissioning, or data center delivery; critical infrastructure and stakeholder leadership experience.
Microsoft Project, Smartsheet, Primavera P6, Jira, Monday.com, BMS, DCIM, EPMS, SCADA, Microsoft Excel
1w
Save
Mark Applied
Hide
Technical Program Manager — New Technology Enabling
Houston, Texas, United States
$60k-$70k/yr OnsiteFull Time
Quest Global
Quest Global: Global engineering services for product development and lifecycle management.
4+ YOEBachelor’s or master’s in engineering or computer science, 4+ years leading engineering programs, and knowledge of PC architecture, firmware, embedded systems, interfaces, vendor coordination, risk management, and stakeholder communication.
ARM, UEFI, USB4, Thunderbolt, USB-C PD, BIOS, Agile, DisplayPort, I²C, UART, embedded systems, PC platform architecture
1w
Save
Mark Applied
Hide
Principal Technical Program Manager - FDE
United States or Redmond or Mountain View or Houston or Austin or Chicago or Cambridge or New York City or Atlanta
$143k-$275k/yr RemoteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
6+ YOEBachelor's degree or equivalent experience; 6+ years in engineering, technical program management, data analysis, or product development; 3+ years managing cross-functional projects; coding and customer-facing experience preferred.
Microsoft Visual Studio Code, GitHub, AI agents
2w
Save
Mark Applied
Hide
Technical Program Manager, Director
Abu Dhabi or Amsterdam or Cambridge or Houston
OnsiteFull Time
Context Labs
Context Labs: Develops data fabric software for trusted climate and ESG reporting.
Requires global enterprise operations, leadership, cloud ecosystems, security, compliance, deployment frameworks, data integration, communication, and complex software implementation experience.
Atlassian