Moveworks
Posted 4d ago

Staff Machine Learning Engineer, Agent Eval Platform

Moveworks
Mountain View, California, United States
OnsiteFull Time
Responsibilities
  • designing judges
  • calibrating evaluations
  • fine-tuning models
Requirements
  • Requires 8+ years in applied ML
  • Data science, or ML engineering
  • Strong Python and applied ML fundamentals
  • Production code experience
  • Expertise in at least three listed evaluation
  • Annotation
  • Ranking
  • Fine-tuning
  • Reward modeling
  • Trajectory analysis, or prompt engineering areas
Technical tools mentioned
PythonLLMSFTRLHFRLAIF

Job description

Company Description

Who we are

Moveworks: the Agentic AI Assistant platform that empowers the entire workforce. 

Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Powered by the world's most advanced LLMs, our proprietary models, and a sophisticated Agentic AI platform, we're transforming how work gets done by allowing AI to take initiative, streamline complex workflows, and continuously learn and adapt.

Moveworks is trusted by over 5.5 million employees at more than 350 of the world’s largest companies, including 10% of the Fortune 500, to automate everyday tasks and streamline business operations. Recognized on the Forbes Cloud 100 and AI 50 lists, Moveworks was also named one of Fast Company’s 2025 Most Innovative Companies and Inc’s Best in Business, in the Best in Innovation category. Moveworks was also recognized at Microsoft’s 2025 Partner of the Year and in 2024, received the AI Breakthrough Award. 

In December 2025, Moveworks was acquired by ServiceNow, marking a pivotal milestone in our journey to create a single front door to work for all business systems. By combining ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform for every person and every workflow. Built to go beyond basic summaries to deliver meaningful business impact. Together, our AI acts across enterprise systems to turn conversations into completed work.

By joining our team, you’ll be at the forefront of the AI transformation, backed by the global scale of ServiceNow and the agility of a high-growth company. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business. Come join us!

ServiceNow: it all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.

Job Description

The Role

Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better?

That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.

This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.

 


What you get to do in this role:

Judge design and calibration

  • A shared base judge with per-item rubrics expressed as configuration next to the dataset — so eval authors express intent, rather than forking a prompt per eval
  • Splitting the problem correctly: deterministic validators for checkable world state ("was the ticket created, with the right item, routed to the right approver?"), and an LLM judge for the parts that are genuinely fuzzy — was the clarifying question appropriate, was policy followed, was the path efficient
  • Scoring that reports its own confidence, so uncertain judgements route to a human instead of quietly becoming training data
  • A standing calibration loop against human-labeled trajectories, run in partnership with our annotation team — they own the human labeling, you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be, so raising that ceiling is part of the job
  • Fine-tuning a small judge model where an off-the-shelf one isn't good enough
  • Guarding against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction
  • Offline↔online divergence: when simulation and production disagree, being the person who can say why, and keeping the suite re-seeded from new production failures so it can't quietly overfit

Self-learning for the agent harness

This is where the pillar is headed, and a large part of why the seat exists.

  • A calibrated trajectory judge is, functionally, a reward model. Turning ours into a process reward model — a dense, step-level signal for what a good agent trajectory looks like — is the unlock
  • Using that signal to optimize the agent itself: prompts, tool selection, planner behavior, retrieval, routing — tuned against simulation rather than against production traffic
  • Building the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement
  • Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal, and dangerously brittle as a definition of success

 

 

Qualifications

To be successful in this role you have:

  • 8+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used
  • Experience turning subjective human judgement into a measurement that holds up — one that other people, and ideally other models, can act on. This is the core of the job
  • Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain
  • Strong Python, and the discipline to ship production-grade code rather than notebooks
  • Ability to think and communicate clearly about complex problems — a large part of this job is convincing engineers that a number means what you say it means, and being right
  • A high degree of ownership and a bias toward shipping at startup pace
  • Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on

Experience in at least 3 of these:

  • LLM-as-judge or automated evaluation design, and calibrating it against human judgement
  • Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint
  • Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly
  • Fine-tuning and evaluating small models: SFT, preference tuning, distillation
  • Reward modeling, RLHF/RLAIF, or process reward models
  • Agent trajectory analysis and step-level fault attribution
  • Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes

 

  

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. 

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact [email protected] for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license. 

About Moveworks

Enterprise AI assistant platform that helps organizations automate employee support, search, and workflows.

Year founded
2016
Employees
710
Organization type
Private
Latest investment
Raised $200.00M Series C (2021) — led by Tiger Global, Alkeon Capital
Headquarters
US

Similar jobs

Machine Learning Engineer roles near Mountain View, California
9h
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Fremont, California, United States
$155k-$200k/yr HybridFull Time
Velo3D
Velo3DNasdaq Capital Market: VELO: Public metal additive manufacturer providing printers, software, and engineering services to aerospace, defense, energy, and industrial customers.
Experience with machine learning pipelines for geometric data and production software engineering; strong Python or C++ skills. Manufacturing, IoT, and industrial automation knowledge preferred.
Python, C++, Flow, Assure
1d
Save
Mark Applied
Hide
Founding Machine Learning Engineer
Mountain View, California, United States
$220k-$300k/yr OnsiteFull Time
Clera
Clera: AI-powered talent agent matching candidates to startup roles.
3+ YOE3–10 years as an ML Engineer, Applied Scientist, or Research Engineer; Python and PyTorch, TensorFlow, or JAX proficiency; ML fundamentals, distributed systems, cloud ML infrastructure, and MLOps experience.
Python, PyTorch, TensorFlow, JAX, AWS, GCP, Azure, Weights & Biases, MLflow
2d
Save
Mark Applied
Hide
Machine Learning Engineer Graduate (E-Commerce Knowledge Graph) - 2027 Start
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
$128k-$317k/yr OnsiteFull Time
TikTok
TikTok: Short-form mobile video and social media platform.
Bachelor's degree in software development, computer science, computer engineering, or related field; machine learning exposure; programming familiarity; strong teamwork and communication skills.
C++, Python, Go, Java, TensorFlow, PyTorch, Hadoop, Spark, Hive, Flink
2d
Save
Mark Applied
Hide
Machine Learning Engineer III
San Jose or Seattle
$146k-$221k/yr HybridFull Time
Expedia Group
Expedia GroupNASDAQ: EXPE: Global travel technology powering online booking platforms.
3+ YOEBachelor’s or master’s degree in a quantitative field, 3+ years of machine learning or data-driven systems experience, Python and ML framework proficiency, and experience with data pipelines and large datasets.
Python, PyTorch, TensorFlow, Spark, SQL, Databricks, AWS
2d
Save
Mark Applied
Hide
Machine Learning Engineer
San Francisco, California, United States
$260k-$300k/yr OnsiteFull Time
Hyperbound
Hyperbound: Private AI sales-coaching platform helping enterprise revenue teams practice conversations and improve performance.
Own machine learning models end to end, including training, fine-tuning, production deployment, on-device execution, evaluation frameworks, benchmarks, and regression suites.
open source models
3d
Save
Mark Applied
Hide
Machine Learning Engineer, Causal Inference, Level 5
Los Angeles or Seattle or Palo Alto or New York City or Bellevue or Santa Monica or California or Washington or New York City
$209k-$313k/yr OnsiteFull Time
Snap Inc.
Snap Inc.NYSE: SNAP: Technology focused on augmented reality and communication.
5+ YOEBachelor’s degree or equivalent practical experience and 5+ years of post-bachelor’s machine learning experience, or equivalent master’s/PhD pathways, with causal inference and experimentation expertise.
Python, pandas, NumPy, scikit-learn, CausalML, CausalM, EconML, DoWhy
3d
Save
Mark Applied
Hide
Machine Learning Engineer, Causal Inference, Level 5
Los Angeles or Seattle or Palo Alto or New York City or Bellevue or Santa Monica or California or Washington or New York
$209k-$313k/yr OnsiteFull Time
Snap Inc.
Snap Inc.NYSE: SNAP: Technology focused on augmented reality and communication.
5+ YOEBachelor’s degree or equivalent experience plus 5+ years in machine learning, or equivalent master’s/PhD pathways; causal inference, experimentation, production modeling, and Python expertise required.
Python, pandas, NumPy, scikit-learn, CausalM, CausalML, EconML, DoWhy
3d
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Sunnyvale, California, United States
$202k-$224k/yr HybridFull Time
Uber Technologies, Inc.
Uber Technologies, Inc.NYSE: UBER: Global technology platform for ride-hailing, delivery, and freight logistics.
4+ YOE4+ years in ML or robotics, bachelor's degree in computer science, computer engineering, or related field, Python and Linux proficiency, and familiarity with modern AI/ML frameworks. Autonomous driving experience preferred.
Python, Linux, PyTorch, C++