1,052 ai evaluation jobs at 564 companies in United States

PromotedHiringCafe
Founding Machine Learning / AI Search Engineer
Cupertino, CA, US
$160k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Build the ML and AI search behind HiringCafe — ranking, recommenders, retrieval, and LLM agents that surface jobs people would never find on their own.
Python, PyTorch, Elasticsearch, LLMs
1w
Save
Mark Applied
Hide
AI Evaluation Lead
United States
$120k-$140k/yr RemoteFull Time
Elly
Elly: AI-native hiring platform that automates recruiting and applicant tracking.
Experience evaluating high-stakes AI/ML outputs, judgment on model behavior, personal finance literacy, familiarity with evaluation and observability tooling, and ability to translate findings into actionable recommendations.
LLM
1mo
Save
Mark Applied
Hide
AI Evaluation Lead
United States
$120k-$140k/yr RemoteFull Time
Fintentional
Fintentional: AI-powered financial planning and personalized investment guidance.
Experience evaluating AI/ML system quality in high-stakes contexts, personal finance literacy, analytical judgment, familiarity with LLM behavior and evaluation/observability tooling, and ability to translate findings into actionable recommendations.
2w
Save
Mark Applied
Hide
AI Evaluation Scientist
McLean, Virginia, United States
$105k-$145k/yr OnsiteFull Time
Steampunk
Steampunk: Provides digital transformation and IT consulting for federal agencies.
2+ YOEDesign and run AI evaluation frameworks, build automated tests and benchmarks, perform LLM/RAG audits, and map findings to responsible AI principles; requires advanced AI evaluation experience and programming proficiency.
Python, PyTorch, Hugging Face, scikit-learn, LangChain, Ragas
2w
Save
Mark Applied
Hide
AI Evaluation Scientist
McLean, Virginia, United States
OnsiteFull Time
Steampunk
Steampunk: Design-led federal technology consulting and innovation services firm.
2+ YOEDesign and run AI evaluation frameworks, build automated evaluation pipelines, create benchmark datasets, perform LLM/RAG error analysis; requires Python and ML library proficiency and 2+ years evaluating ML or NLP models.
Python, PyTorch, Hugging Face, scikit-learn, LangChain, Ragas, OWASP LLM Top 10
1w
Save
Mark Applied
Hide
Director AI Evaluation
Pennsylvania, United States
RemoteFull Time
Geisinger
Geisinger: Provides medical care, insurance plans, and medical education.
8+ YOE3+ MgmtLead AI evaluation standards and teams; hands-on model evaluation and monitoring; 8+ years experience with 3+ years managerial, Python and SQL fluency, experimental design and fairness expertise.
Python, SQL, LLM
2w
Save
Mark Applied
Hide
AI Evaluation Scientist
McLean, Virginia, United States
$105k-$145k/yr OnsiteFull Time
Steampunk
Steampunk: Federal digital transformation and consulting services provider.
2+ YOEAbility to design and run AI evaluation frameworks, build automated evaluation pipelines, analyze LLM/RAG behavior, and communicate findings. Requires public trust eligibility, 2+ years evaluating ML/NLP/LLM systems, and proficiency in Python and ML libraries.
Python, PyTorch, Hugging Face, scikit-learn, LangChain, Ragas, OWASP LLM Top 10
1mo
Save
Mark Applied
Hide
AI Evaluation Engineer
Bucharest or Boulder or Boston or Chicago or Glendora or Richmond or London or Auckland
lei16k-lei19k/mo HybridFull Time
Yes Energy
Yes Energy: Provides real-time electric power market data and trading analytics.
5+ YOEBachelor's in CS/Data/Math, 5+ years software engineering with focus on testing/evaluation or AI/ML tooling, strong Python, experience building evaluation pipelines and LLM eval concepts, familiarity with RAG and production data pipelines.
RAGAS, LangSmith, PromptFlow, Snowflake, Python, RAG
2mo
Save
Mark Applied
Hide
Clinical AI Evaluation Specialist
United States
$90k-$115k/yr RemoteFull Time
CentralReach
CentralReach: Software for autism care and special education therapy providers
Experience in QA/evaluation of healthcare AI, governance frameworks, and collaboration with product/engineering.
2w
Save
Mark Applied
Hide
Senior Director - AI Evaluation Platform
Santa Clara, California, United States
$279k-$488k/yr HybridFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
Proven leadership of large AI or platform engineering teams, deep expertise in AI evaluation methodologies, production evaluation platforms, strong communication, and hands-on experience with PyTorch or TensorFlow.
PyTorch, TensorFlow
1d
Save
Mark Applied
Hide
AI Enablement & Governance– AI Quality & Evaluation Lead
Illinois, United States
$150k-$200k/yr RemoteFull Time
Alight
AlightNYSE: ALIT: Provides cloud-based HR, payroll, and benefits administration services.
5+ YOE5+ years in data science/ML engineering or AI quality, expertise in evaluation, statistical validation, system design for RAG/Agents, and stakeholder enablement.
Python, Pandas, Scikit-learn, RAGAS, TruLens, MLflow
1mo
Save
Mark Applied
Hide
AI Evaluation Subject Matter Expert
Charleston, South Carolina, United States
HybridFull Time
Foxhole Technology
Foxhole Technology: Cybersecurity and IT services provider delivering mission-focused solutions to federal civilian and defense agencies.
Bachelor’s or higher in a technical field; extensive AI evaluation/test experience; active DoD Secret clearance with TS capability; strong communication.
Python, R, SQL, Jupyter, Machine Learning frameworks
1mo
Save
Mark Applied
Hide
Speech AI Evaluation Specialist - Vietnamese (USA)
United States
$15/hr RemotePart Time, Contract
RWS
RWSLondon Stock Exchange: RWS: Provides AI training data, translation, and intellectual property services.
Native-level Vietnamese, fluent English (B2–C2), participate in short voice conversations, evaluate AI responses and provide objective ratings; flexible part-time freelance work, AI/data experience preferred.
2mo
Save
Mark Applied
Hide
MD - Principal Clinical AI Evaluation Strategist
United States
$95k-$159k/yr RemoteFull Time
RELX
RELXLondon Stock Exchange: REL: Provides information-based analytics and decision tools for professional customers.
10+ YOEMD with 10+ years clinical experience; experience evaluating AI in clinical applications; leading large expert teams; medical informatics expertise; medical content validation.
2mo
Save
Mark Applied
Hide
MD - Principal Clinical AI Evaluation Strategist
United States
$95k-$159k/yr RemoteFull Time
Elsevier
ElsevierLondon Stock Exchange: REL: Provides scientific, technical, and medical research information and analytics.
10+ YOEMD with 10+ years clinical experience and medical informatics expertise; experience evaluating AI in clinical applications; proven leadership of expert teams; strong medical knowledge systems background.
1w
Save
Mark Applied
Hide
Sr. Manager, Test Automation & AI Evaluation
Newark, New Jersey, United States
$98k-$135k/yr OnsiteFull Time
WebMD
WebMD: Providing digital health information services for consumers and healthcare professionals.
7+ YOE7+ years in quality engineering/test automation with leadership; hands-on Playwright/Cypress/Selenium; CI/CD integration; experience evaluating AI/LLM systems and defining test strategy.
Playwright, Cypress, Selenium
1mo
Save
Mark Applied
Hide
AI Engineer, Evaluation
San Francisco or New York City
$150k-$250k/yr HybridFull Time
Distyl AI
Distyl AI: Builds AI-native orchestration platforms for enterprise operations.
2+ YOE2+ years software engineering, strong Python, experience with evaluation- or experiment-driven development, ability to encode human judgment into tests/graders, systems-oriented mindset, and willingness to travel 10–50%.
Python, LLM
3w
Save
Mark Applied
Hide
AL/ML Evaluation Engineer
Atlanta, Georgia, United States
$129k-$292k/yr HybridFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Consulting and technology services for government and commercial clients
5+ YOE5+ years in generative AI/LLMs/AI agents, 3+ years AI evaluation in enterprise, Bachelor's in CS/Engineering/Data Science required, Python, PySpark, TensorFlow/PyTorch, MLflow, Palantir Foundry, Azure experience, ability to obtain Public Trust.
Python, PySpark, Palantir Foundry, Foundry AIP, Codex, Claude, TensorFlow, PyTorch, SQL, MLflow, Azure, Jira, CI/CD
3w
Save
Mark Applied
Hide
Lead, Search & Evaluation, AI and Automation Drug Discovery
Cambridge, Massachusetts, United States
$177k-$278k/yr HybridFull Time
Takeda
TakedaTokyo Stock Exchange: 4502: Develops and manufactures pharmaceutical products for global healthcare needs.
Bachelor's in a scientific/technical field required; significant pharma/biotech experience; deep knowledge of drug discovery and AI/ML; experience evaluating external technologies and leading cross-functional diligence; strong communication and stakeholder influence.
AI/ML
1w
Save
Mark Applied
Hide
Senior Investment Banking Subject Matter Expert (AI Evaluation) | U.S.
United States or North America
$55-$60/hr RemoteContract
Volga Partners
Volga Partners: Provides data annotation and language services for AI companies.
Expert-level investment banking and corporate finance experience with advanced financial modeling and valuation skills; strong analytical, written English, and attention to detail; ability to evaluate AI-generated financial analyses.
Bloomberg Terminal, Capital IQ, FactSet, Refinitiv, PitchBook
3w
Save
Mark Applied
Hide
AL/ML Evaluation Engineer
Atlanta, Georgia, United States
$129k-$292k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Provides technology and management consulting services to diverse organizations.
5+ YOE5+ years Generative AI/LLM experience, 3+ years AI agents evaluation, 2+ years deep research evaluation; proficiency in Python, TensorFlow/PyTorch, PySpark, Palantir Foundry, MLflow, Microsoft Azure; Bachelor’s in CS/Engineering/Data Science; ability to obtain Public Trust.
PySpark, Palantir Foundry, Foundry AIP, Codex, Claude, TensorFlow, PyTorch, SQL, MLflow, Microsoft Azure, Jira