275 evaluations engineer jobs at 170 companies in Fairfield, CA
1mo
Save
Mark Applied
Hide
1mo
Evaluations Engineer
San Francisco, California, United States
$130k-$165k/yrOnsiteFull Time
Vals AI: Building enterprise benchmarks for evaluating LLM performance.
Strong engineering fundamentals, professional Python expertise, familiarity with LLMs, Git workflow experience, ability to analyze model error modes and work with cross-functional teams. In-person in San Francisco; relocation/transportation support provided.
Evaluations Engineering - Member of Technical Staff
San Francisco, California, United States
$200k-$400k/yrOnsiteFull Time
Simile: Building AI simulations to predict human behavior at scale.
Several years building production backend and data systems, experience with evaluation pipelines, ML/LLM fluency, strong system design and communication skills.
Experience building secure, isolated infrastructure for sensitive model evaluations; strong Python skills; security engineering fundamentals; data pipeline and sandboxing experience.
Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: Develops AI-powered voice recorders and automated transcription software.
Python software engineering; building distributed systems, data pipelines, and evaluation harnesses at scale; partner with ML researchers to define benchmarks; build dashboards and monitor model health; debug mid-training anomalies; communicate results clearly.
Waymo: Autonomous driving technology for ride-hailing and logistics.
4+ YOEBS in a quantitative field; 4-7 years experience; strong C++, Python, SQL skills; data analysis; building data pipelines; ML familiarity; strong coding standards.
C++, Python, SQL, Machine Learning, Data Processing, Statistics
Distyl AI: Builds AI-native orchestration platforms for enterprise operations.
2+ YOE2+ years software engineering, strong Python, experience with evaluation- or experiment-driven development, ability to encode human judgment into tests/graders, systems-oriented mindset, and willingness to travel 10–50%.
Ambral: AI-powered account managers for enterprise customer success teams.
4+ YOE4+ years building production software or ML systems, including 2+ years with RL environments or evaluation/agent infrastructure; strong software engineering, data processing, reproducibility, and evaluation skills.
Granica: Provides an AI efficiency platform for optimizing enterprise training data.
5+ YOERequires 5+ years in sales engineering, solutions architecture, forward deployed or data engineering; expertise in data platforms, distributed systems, cloud infrastructure, SQL, and programming; enterprise evaluations and deployments experience.
Apache Spark, Apache Iceberg, Delta Lake, Trino, Presto, SQL, Large Tabular Models (LTMs)
Mira Mace: AI-powered healthcare advocacy and navigation for Medicare beneficiaries.
Experience building and shipping LLM-powered systems or agents, strong AI fluency (prompting, retrieval, evaluation), production monitoring, and solid software engineering practices.
Bunkerhill Health: AI-powered platform automating clinical reasoning and diagnostic workflows.
Strong engineering skills, clinical stakeholder communication, production debugging, AI agent development, evaluation building, systems integration, and ownership of customer outcomes are required.
MaintainX: Provides mobile-first software for maintenance and asset management.
Applied GenAI experience with prompt engineering, structured outputs, RAG, evaluation datasets, production LLM features, multimodal inputs, and model quality, cost, and latency optimization.
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
4+ YOEBachelor's, master's, or PhD in imaging science, electrical engineering, computer science, or similar; 4+ years of camera ISP experience; ISP tuning, sensor characterization, image quality evaluation, and C/C++ or Python proficiency.
Magic Patterns: AI-powered platform for rapid software prototyping and design.
Experience using AI tools, working with AI models or agents, and building evaluation systems, post-training, or context engineering; strong first-principles problem-solving skills.
Clera: AI talent agent matching professionals with high-growth startup roles
2+ YOE2–4 years in applied research or forward-deployed engineering; Python, Docker, and Linux proficiency; benchmark and evaluation experience; strong debugging, communication, and independent problem-solving skills.
Bardeen: AI-powered platform for automating browser-based workflows and tasks.
1+ YOEMaster's in computational science and engineering or related machine learning field plus 1 year of experience with deep learning, LLMs, Python, cloud infrastructure, evaluation, monitoring, and distributed inference.
OpenAI: Develops artificial intelligence models and generative AI software services.
4+ YOE4+ years software engineering with customer-facing delivery; strong prototyping, production ownership; proficiency in Python and/or JavaScript/TypeScript; cloud deployment experience; practical AI/LLM and evaluation experience.
Abby Care: Trains and pays family members to provide home care.
2+ YOE2+ years professional software/ML or applied AI experience, strong software engineering, experience with language models/AI APIs, backend and data pipeline skills, ability to evaluate models and work with healthcare workflows.
Artos: AI-powered document authoring platform for life sciences R&D.
2+ YOE2+ years building and deploying AI/ML applications, hands-on LLM experience, backend engineering with Python, API development (FastAPI/Django), cloud container deployment, R&D on model capabilities, and evaluation tooling experience.