Technical Program Manager, Evaluation & Validation
Sunnyvale, California, United States
$215k-$274k/yrHybridFull Time
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
8+ YOE8+ years leading complex multi-team programs; technical credibility in simulation, evaluation, ML; strong program management, communication, data-driven decision-making, and systems thinking.
Lapel: Unifying customer data to power personalized service at scale.
3+ YOE3+ years recruiting across technical and non-technical roles; owned end-to-end hiring; strong sourcing, candidate evaluation, and stakeholder partnership skills; comfortable working with founders and moving quickly.
Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Senior SDET to lead design and implementation of automated model evaluation, build LLM evaluation pipelines and infrastructure, and partner with modeling, framework, and infrastructure teams.
Stripe: Provides online payment processing and financial infrastructure for businesses.
5+ YOEPhD or MD in immunology/infectious disease with 5+ years post-grad research; portfolio management, grant/investment experience; product development and project management; strong scientific evaluation and communication skills.
Squint: AI and AR platform for frontline manufacturing workers.
2+ YOE2+ years in operations/strategy/TPM or applied AI, hands-on experience building and shipping AI tools or agents, familiarity with agent tooling, light coding ability, strong communication and evaluation skills.
TencentHong Kong Stock Exchange: 0700: Global provider of internet services and digital technology solutions.
Design end-to-end data acquisition strategies, manage vendor relationships, own data pipelines, set quality/evaluation metrics, and navigate data governance (including CCPA/CPRA) for AI research initiatives.
Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP)
McLean or San Francisco or New York City or San Jose or Cambridge
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
4+ YOEBachelor's plus 6 years or master's plus 4 years in AI/ML development; 6 years programming in Python, Go, Scala, or Java; experience building scalable AI systems and leading engineering teams.
Staff/Lead Data Scientist (Model Evaluation) - TikTok Integrity and Safety (San Jose)
San Jose, California, United States
$219k-$422k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
5+ YOE5+ years data science experience, Bachelor's degree, strong SQL and Python/R skills, visualization, ML knowledge, causal inference and A/B testing experience, experiment evaluation and mentoring experience.
Technical Program Manager, Evaluation & Validation
Sunnyvale or London
$215k-$274k/yrHybridFull Time
Wayve: Develops AI software for autonomous vehicle navigation.
8+ YOE8+ years leading complex multi-team programs, strong technical credibility with simulation/ML teams, program and risk management, data-driven decision making, clear communication and systems thinking.
8+ YOEBachelor's degree or equivalent practical experience; 8 years in technical, engineering, research, policy, or regulatory roles; experience with foundational model lifecycles, training data, evaluation, deployment, and cross-functional teams.
AccentureNYSE: ACN: Global provider of management consulting and technology services.
5+ YOEMinimum 5 years' experience designing and delivering training curricula, onboarding and competency validation, training documentation/versioning, data-driven evaluation of training effectiveness, and coaching staff.
OpenAI: Develops artificial intelligence models and generative AI software services.
8+ YOERequires 8+ years in infrastructure strategy, site selection, energy, data centers, economic development, corporate development, international expansion, or related work, plus market evaluation and cross-functional coordination experience.
Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP)
McLean or San Francisco or Cambridge or San Jose or New York City
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's plus 6+ years (or Master's plus 4+ years) developing AI/ML systems; strong programming (Python, Go, Scala, Java); cloud AI deployment experience; leadership and research application skills.
Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP)
McLean or San Francisco or New York City or San Jose or Cambridge
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years in AI/ML; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment and team leadership preferred.
Staff Technical Lead Manager, Prediction & Planning, ML Eval
Mountain View or San Francisco
$251k-$310k/yrHybridFull Time
Waymo: Autonomous driving technology for ride-hailing and logistics.
5+ YOE3+ MgmtMS in CS/Math or equivalent; 5+ years distributed infra; 3+ years engineering management; strong Python/C++; ML evaluation knowledge; familiarity with ML deployment/ orchestration tools.
Anthropic: Developing safe and reliable artificial intelligence systems.
Lead distributed-systems infrastructure for model evaluation; 1+ years managing engineers or tech-lead-with-reports; strong Python and Rust; experience with high-throughput fault-tolerant systems; bachelor\u0002s or equivalent.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOEBS/MS in EE/CE/CS/Systems or equivalent; 12+ years in product performance and power evaluations; experience with system-level features, product binning, data analysis/statistics, NPI, critical path and power/performance analysis, silicon bringup/validation, and offshore collaboration.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years in data science/model evaluation/experimentation, Master's/PhD or equivalent, proven model evaluation for go/no-go decisions, hands-on Python and statistical/experimental methods, LLM/GenAI benchmarking experience.