27 inference engineer jobs at 9 companies in Fredericksburg, VA

1d
Save
Mark Applied
Hide
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
FSDP, DeepSpeed, Megatron, NCCL, NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, Python, C++, CUDA, Triton, CUTLASS, Nsight Systems, Nsight Compute, PyTorch Profiler, perf, Kubernetes, Slurm, ROS, OpenCV
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York City or McLean or San Jose or Cambridge
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor’s plus 6 years or master’s plus 4 years in AI/ML development; 6 years programming in Python, Go, Scala, or Java; cloud AI deployment and engineering leadership preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
1w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
1mo
Save
Mark Applied
Hide
Staff AI Infrastructure Engineer
Austin or Reston
HybridFull Time
Seekr
Seekr: Transparent AI platform for enterprise and government decision-making.
8+ YOE8+ years building distributed systems and cloud-native AI infrastructure; strong Python and systems programming skills; Kubernetes, GPU inference, and platform engineering experience; leadership and architecture experience.
Python, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, Infrastructure-as-Code, GitOps, CI/CD, AWS, Azure, Oracle Cloud Infrastructure, Google Cloud Platform
1mo
Save
Mark Applied
Hide
Senior AI/ML Engineer
United States or Minneapolis or Washington
$120k-$215k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
3+ YOEBachelor's in CS or quantitative field; 3+ years programming (Python, SQL, packaging), experience with observability, model registry/experiment tracking, responsible AI, GenAI/agents, serving/inference, and cloud platforms.
Python, pytest, SQL, shell, LangChain, Vertex AI, MLflow, AWS, Azure, GPUs, TPUs
1w
Save
Mark Applied
Hide
ML Ops Engineer
Chantilly, Virginia, United States
HybridFull Time
Bana Solutions
Bana Solutions: Provides software engineering, cybersecurity, and data solutions for government agencies.
6+ YOEOwnership of full production ML lifecycle, model serving for low-latency inference, monitoring/drift detection, automated retraining pipelines, Kubernetes/Docker deployment, strong Python and software engineering skills.
MLflow, SageMaker, Weights & Biases, Vertex, PyTorch, TensorFlow, ONNX, Feast, Kubernetes, Docker, Helm, Prometheus, Grafana, Evidently, TorchServe, Triton, KServe, Seldon, BentoML, Apache Airflow, Kafka, Bytewax, Flink, Spark Streaming, Kafka Streams, CUDA, NVIDIA, S3, MinIO, Redshift, EKS, OpenShift, Python
4w
Save
Mark Applied
Hide
AI Integration Engineer
Annapolis Junction or McLean
$113k-$257k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Provides technology and management consulting services to diverse organizations.
3+ YOE3+ years building production software in Python, integrating services/APIs and LLMs, implementing inference flows, writing integration tests, and knowledge of HTTP auth patterns and container/CI/CD tooling.
Python, OAuth2, JWT, Kafka, SQS, SNS, AWS Bedrock AgentCore, Google Gemini Enterprise Agent Platform, Microsoft Foundry Agent Service, Docker, Kubernetes, Redis, Postgres, LangGraph, CrewAI, AutoGen
1mo
Save
Mark Applied
Hide
Sr. GenAI Specialist PSA , Solutions Architecture
Herndon, Virginia, United States
$154k-$208k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE7+ years designing and operating distributed applications; 5+ years customer-facing experience building production GenAI/ML systems; hands-on LLM, RAG, agentic workflows, prompt engineering, inference optimization; experience with AWS AI/ML services.
Amazon Bedrock, SageMaker, AgentCore, Bedrock Agents, Strands Agents SDK, LangChain, OpenSearch Serverless, pgvector, Pinecone
2w
Save
Mark Applied
Hide
Principal Data Scientist
Glen Allen, Virginia, United States
$200k-$300k/yr OnsiteFull Time
W. R. Berkley
W. R. BerkleyNYSE: WRB: Provides commercial property, casualty, and specialty insurance and reinsurance.
10+ YOE10+ years building and shipping production ML/AI systems; expert Python and ML engineering; experience with LLMs, MLOps, causal inference, and statistical methods; Bachelor's in quantitative field required; advanced degree preferred.
Python, NumPy, Pandas, Scikit-learn, PyTorch, TensorFlow, Hugging Face, LangChain, LlamaIndex, MLflow, Docker, Kubernetes, SQL, dbt, Spark, Databricks, Azure ML, AWS SageMaker, GCP Vertex AI, SHAP, LIME, Git