674 ai evaluation engineer jobs at 412 companies in United States

2mo
Save
Mark Applied
Hide
AI Evaluation Engineer
Bucharest or Boulder or Boston or Chicago or Glendora or Richmond or London or Auckland
lei16k-lei19k/mo HybridFull Time
Yes Energy
Yes Energy: Provides real-time electric power market data and trading analytics.
5+ YOEBachelor's in CS/Data/Math, 5+ years software engineering with focus on testing/evaluation or AI/ML tooling, strong Python, experience building evaluation pipelines and LLM eval concepts, familiarity with RAG and production data pipelines.
RAGAS, LangSmith, PromptFlow, Snowflake, Python, RAG
2d
Save
Mark Applied
Hide
AI Evaluation Engineer (QA)
New York City or Austin or Miami or Dallas or Phoenix or São Paulo or Toronto
OnsiteFull Time, Contract
Appnovation
Appnovation: Digital consultancy providing strategy, design, and software engineering services.
4+ YOEBachelor's degree in a technical field or equivalent experience, 4+ years in QA/test engineering, Python and data-science skills, LLM evaluation and statistical analysis experience, and automated test framework expertise.
Python, LLM evaluation frameworks, CI/CD, Test automation frameworks
2mo
Save
Mark Applied
Hide
AI Engineer, Evaluation
San Francisco or New York City
$150k-$250k/yr HybridFull Time
Distyl AI
Distyl AI: Builds AI-native orchestration platforms for enterprise operations.
2+ YOE2+ years software engineering, strong Python, experience with evaluation- or experiment-driven development, ability to encode human judgment into tests/graders, systems-oriented mindset, and willingness to travel 10–50%.
Python, LLM
2w
Save
Mark Applied
Hide
AI Engineer – Algorithm Evaluation & Agentic Systems
Sunnyvale, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience evaluating AI algorithms and designing agentic systems, with system-level reasoning, metric definition, engineering rigor, and clear technical communication.
1mo
Save
Mark Applied
Hide
Senior Director - AI Evaluation Platform
Santa Clara, California, United States
$279k-$488k/yr HybridFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
Proven leadership of large AI or platform engineering teams, deep expertise in AI evaluation methodologies, production evaluation platforms, strong communication, and hands-on experience with PyTorch or TensorFlow.
PyTorch, TensorFlow
1w
Save
Mark Applied
Hide
Sr. Evaluation Engineer
San Francisco, California, United States
$158k-$218k/yr HybridFull Time
LogicMonitor
LogicMonitor: Provides AI-powered hybrid observability for enterprise IT infrastructure.
5+ YOE5+ years in software engineering, machine learning, or applied AI; strong Python skills; experience with AI evaluation, LLMs, agents, retrieval-augmented generation, testing, monitoring, and CI/CD.
Python, LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, CI/CD
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP)
McLean or San Francisco or New York City or San Jose or Cambridge
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
4+ YOEBachelor's plus 6 years or master's plus 4 years in AI/ML development; 6 years programming in Python, Go, Scala, or Java; experience building scalable AI systems and leading engineering teams.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, Google Cloud, Azure, C++, C#, Golang
1mo
Save
Mark Applied
Hide
AI Engineer, Agents and Evaluation
Torrance, California, United States
OnsiteFull Time
Pronoia Energy
Pronoia Energy: Developing room-temperature macroscopic quantum energy storage technology.
Proven production experience shipping agent or tool-use systems, experience self-hosting open-weights models on GPUs, strong orchestration/serving/evaluation skills, and habit of treating hallucination as engineering problem.
vLLM, TGI
1w
Save
Mark Applied
Hide
Observability & Evaluation Engineer
Charlotte, North Carolina, United States
HybridFull Time
NTT DATA
NTT DATATokyo Stock Exchange: 9613: Provides global digital technology consulting and IT services.
7+ YOERequires 7+ years of engineering experience, 5+ years of Python, observability, LLM evaluation, Agile, and production support experience; knowledge of RAG, AI quality, telemetry, and monitoring.
Python, Kubernetes, CI/CD
1mo
Save
Mark Applied
Hide
Senior AI Quality Engineer (LLM Evaluation & Automation) 1754
United States
RemoteFull Time
Softgic
Softgic: Custom software development and digital transformation technology consulting firm.
Experience evaluating ML/LLM systems, strong test and benchmark design, comfort with noisy/probabilistic metrics, and scripting/automation skills.
AI, LLM, ML, CI
1mo
Save
Mark Applied
Hide
Senior Applied AI Engineer (GenAI & LLM Evaluation)
United Kingdom or United States
RemoteFull Time
Ascent
Ascent: Delivers AI-driven digital transformation and software engineering services.
5+ YOE5+ years in data engineering/ML/AI engineering, strong Python and SQL, hands-on Generative AI and LLM evaluation, building production ML solutions, data pipelines, validation, and API/cloud integration.
Python, SQL, Large Language Models (LLMs)
1w
Save
Mark Applied
Hide
Software Engineer, AI Evaluation
San Francisco, California, United States
$148k-$232k/yr HybridFull Time
Nuna
Nuna: Provides data analytics and software for value-based healthcare solutions.
Significant production systems experience; expertise evaluating AI systems, agentic tooling, adversarial testing, statistics, experimental design, and functional UI development for non-engineers.
LangSmith, Braintrust, DeepEval, Ragas, Promptfoo
2mo
Save
Mark Applied
Hide
Principal Test & Evaluation Engineer
Melbourne, Florida, United States
$89k-$134k/yr OnsiteFull Time
Northrop Grumman
Northrop GrummanNYSE: NOC: Designs and manufactures advanced aerospace and defense systems.
1+ YOESTEM degree (BS+5yrs | MS+3yrs | PhD+1yr), active US Secret clearance (with ability to obtain SAP), experience with virtualization/containerization, on-premise and cloud (Azure/AWS), Python scripting on Linux/Windows, Agile, demo/experiment integration, ability to lift 75 lbs, willingness to travel up to 25%.
Python, Linux, Windows, Azure, AWS, Kubernetes, O-RAN, 5G, AI/ML, Model-Based Systems Engineering (MBSE), Agile, microservices
2mo
Save
Mark Applied
Hide
Principal Test & Evaluation Engineer
Melbourne or United States
$89k-$134k/yr OnsiteFull Time
Northrop Grumman
Northrop GrummanNYSE: NOC: Develops and manufactures advanced aerospace, defense, and space systems.
1+ YOESTEM degree (BS/MS/PhD) with 5/3/1 years respectively, active U.S. Secret clearance, experience with 5G/O-RAN, cloud (Azure/AWS), virtualization/containerization, Python on Linux/Windows, Agile, ability to lift 75 lbs, willingness to travel up to 25%.
Python, Linux, Windows, Azure, AWS, Kubernetes, O-RAN, Model-Based Systems Engineering (MBSE), AI/ML, Agile
2mo
Save
Mark Applied
Hide
AI Engineer - SME
Annapolis Junction, Maryland, United States
$206k-$221k/yr OnsiteFull Time
Gormat
Gormat: Provides cybersecurity and systems engineering services to government defense agencies
Design, build, integrate, or evaluate AI/ML-enabled applications; experience with LLMs, agentic systems, and AI frameworks; strong communication.
LangChain, LangGraph, Semantic Kernel, AutoGen, CrewAI, Haystack, LlamaIndex, Python, JavaScript, TypeScript
1mo
Save
Mark Applied
Hide
Staff AI Engineer
United States
$250k-$265k/yr RemoteFull Time
Twin Health
Twin Health: Improving metabolic health through personalized AI digital twin technology
8+ YOE8+ years building AI/ML systems in production; deep software engineering fundamentals; experience with LLMs, Generative AI, RAG, evaluation frameworks, CI/CD, and Python; Bachelor's or Master's in CS/Statistics or related; U.S. work authorization required.
Python, LLMs, Generative AI, RAG, Prompt Engineering, CI/CD
1w
Save
Mark Applied
Hide
Applied AI Lead - Evaluation & Measurement
Frederick or Beavercreek or Fort Walton Beach or Washington
OnsiteFull Time
Leonardo DRS
Leonardo DRSNASDAQ: DRS: Manufactures advanced electronic systems for defense and military applications.
5+ YOERequires 5+ years in data, ML, or test engineering, including 3+ years in evaluation or model validation and recent hands-on LLM evaluation. Requires a bachelor's degree or equivalent, U.S. citizenship, and active or obtainable DOD clearance.
pipelines, dashboards, large-language-model, telemetry
1mo
Save
Mark Applied
Hide
Principal AI Engineer
United States or Pennsylvania
$160k-$208k/yr RemoteFull Time
Vertex
VertexNASDAQ: VERX: Automates global indirect tax calculation and compliance for enterprises.
12+ YOE12+ years software/AI engineering with hands-on experience building LLM orchestration, MCP servers, agents, and retrieval/RAG systems; bachelor's in CS or related; strong evaluation, observability, and stakeholder skills.
LangGraph, LlamaIndex, Semantic Kernel, Model Context Protocol (MCP)
6d
Save
Mark Applied
Hide
Applied AI Engineer
San Francisco, California, United States
OnsiteFull Time
Magic Patterns
Magic Patterns: AI-powered platform for rapid software prototyping and design.
Experience using AI tools, working with AI models or agents, and building evaluation systems, post-training, or context engineering; strong first-principles problem-solving skills.
3w
Save
Mark Applied
Hide
AI Solutions Engineer
Parsippany or Nashville
$140k-$170k/yr HybridFull Time
Managed Health Care Associates
Managed Health Care Associates: Provides group purchasing and software for post-acute care providers.
5+ YOE5+ years software engineering experience; deep experience with agentic AI, prompt engineering, agent orchestration, CI/CD, Azure/M365, AI security, and evaluation frameworks; strong communication and cloud-native skills.
GitHub, Visual Studio Code, GitHub Copilot, Claude, Azure, M365, Azure DevOps, GitHub Actions, SonarQube, Playwright, Jira, Azure Monitor, Application Insights, Model Context Protocol (MCP)