1,052 ai evaluation jobs at 564 companies in United States
🚀PromotedHiringCafe
Founding Machine Learning / AI Search Engineer
Cupertino, CA, US
$160k-$310k/yrOn-SiteFull Time
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Build the ML and AI search behind HiringCafe — ranking, recommenders, retrieval, and LLM agents that surface jobs people would never find on their own.
Elly: AI-native hiring platform that automates recruiting and applicant tracking.
Experience evaluating high-stakes AI/ML outputs, judgment on model behavior, personal finance literacy, familiarity with evaluation and observability tooling, and ability to translate findings into actionable recommendations.
Fintentional: AI-powered financial planning and personalized investment guidance.
Experience evaluating AI/ML system quality in high-stakes contexts, personal finance literacy, analytical judgment, familiarity with LLM behavior and evaluation/observability tooling, and ability to translate findings into actionable recommendations.
Steampunk: Provides digital transformation and IT consulting for federal agencies.
2+ YOEDesign and run AI evaluation frameworks, build automated tests and benchmarks, perform LLM/RAG audits, and map findings to responsible AI principles; requires advanced AI evaluation experience and programming proficiency.
Steampunk: Design-led federal technology consulting and innovation services firm.
2+ YOEDesign and run AI evaluation frameworks, build automated evaluation pipelines, create benchmark datasets, perform LLM/RAG error analysis; requires Python and ML library proficiency and 2+ years evaluating ML or NLP models.
Geisinger: Provides medical care, insurance plans, and medical education.
8+ YOE3+ MgmtLead AI evaluation standards and teams; hands-on model evaluation and monitoring; 8+ years experience with 3+ years managerial, Python and SQL fluency, experimental design and fairness expertise.
Steampunk: Federal digital transformation and consulting services provider.
2+ YOEAbility to design and run AI evaluation frameworks, build automated evaluation pipelines, analyze LLM/RAG behavior, and communicate findings. Requires public trust eligibility, 2+ years evaluating ML/NLP/LLM systems, and proficiency in Python and ML libraries.
Bucharest or Boulder or Boston or Chicago or Glendora or Richmond or London or Auckland
lei16k-lei19k/moHybridFull Time
Yes Energy: Provides real-time electric power market data and trading analytics.
5+ YOEBachelor's in CS/Data/Math, 5+ years software engineering with focus on testing/evaluation or AI/ML tooling, strong Python, experience building evaluation pipelines and LLM eval concepts, familiarity with RAG and production data pipelines.
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
Proven leadership of large AI or platform engineering teams, deep expertise in AI evaluation methodologies, production evaluation platforms, strong communication, and hands-on experience with PyTorch or TensorFlow.
AI Enablement & Governance– AI Quality & Evaluation Lead
Illinois, United States
$150k-$200k/yrRemoteFull Time
AlightNYSE: ALIT: Provides cloud-based HR, payroll, and benefits administration services.
5+ YOE5+ years in data science/ML engineering or AI quality, expertise in evaluation, statistical validation, system design for RAG/Agents, and stakeholder enablement.
Foxhole Technology: Cybersecurity and IT services provider delivering mission-focused solutions to federal civilian and defense agencies.
Bachelor’s or higher in a technical field; extensive AI evaluation/test experience; active DoD Secret clearance with TS capability; strong communication.
Speech AI Evaluation Specialist - Vietnamese (USA)
United States
$15/hrRemotePart Time, Contract
RWSLondon Stock Exchange: RWS: Provides AI training data, translation, and intellectual property services.
Native-level Vietnamese, fluent English (B2–C2), participate in short voice conversations, evaluate AI responses and provide objective ratings; flexible part-time freelance work, AI/data experience preferred.
RELXLondon Stock Exchange: REL: Provides information-based analytics and decision tools for professional customers.
10+ YOEMD with 10+ years clinical experience; experience evaluating AI in clinical applications; leading large expert teams; medical informatics expertise; medical content validation.
ElsevierLondon Stock Exchange: REL: Provides scientific, technical, and medical research information and analytics.
10+ YOEMD with 10+ years clinical experience and medical informatics expertise; experience evaluating AI in clinical applications; proven leadership of expert teams; strong medical knowledge systems background.
WebMD: Providing digital health information services for consumers and healthcare professionals.
7+ YOE7+ years in quality engineering/test automation with leadership; hands-on Playwright/Cypress/Selenium; CI/CD integration; experience evaluating AI/LLM systems and defining test strategy.
Distyl AI: Builds AI-native orchestration platforms for enterprise operations.
2+ YOE2+ years software engineering, strong Python, experience with evaluation- or experiment-driven development, ability to encode human judgment into tests/graders, systems-oriented mindset, and willingness to travel 10–50%.
Booz Allen HamiltonNYSE: BAH: Consulting and technology services for government and commercial clients
5+ YOE5+ years in generative AI/LLMs/AI agents, 3+ years AI evaluation in enterprise, Bachelor's in CS/Engineering/Data Science required, Python, PySpark, TensorFlow/PyTorch, MLflow, Palantir Foundry, Azure experience, ability to obtain Public Trust.
Lead, Search & Evaluation, AI and Automation Drug Discovery
Cambridge, Massachusetts, United States
$177k-$278k/yrHybridFull Time
TakedaTokyo Stock Exchange: 4502: Develops and manufactures pharmaceutical products for global healthcare needs.
Bachelor's in a scientific/technical field required; significant pharma/biotech experience; deep knowledge of drug discovery and AI/ML; experience evaluating external technologies and leading cross-functional diligence; strong communication and stakeholder influence.
Senior Investment Banking Subject Matter Expert (AI Evaluation) | U.S.
United States or North America
$55-$60/hrRemoteContract
Volga Partners: Provides data annotation and language services for AI companies.
Expert-level investment banking and corporate finance experience with advanced financial modeling and valuation skills; strong analytical, written English, and attention to detail; ability to evaluate AI-generated financial analyses.
Bloomberg Terminal, Capital IQ, FactSet, Refinitiv, PitchBook
Booz Allen HamiltonNYSE: BAH: Provides technology and management consulting services to diverse organizations.
5+ YOE5+ years Generative AI/LLM experience, 3+ years AI agents evaluation, 2+ years deep research evaluation; proficiency in Python, TensorFlow/PyTorch, PySpark, Palantir Foundry, MLflow, Microsoft Azure; Bachelor’s in CS/Engineering/Data Science; ability to obtain Public Trust.