36 inference engineer jobs at 13 companies in Arlington, VA

6d
Save
Mark Applied
Hide
Software Engineer - Inference
Laurel, Maryland, United States
$85k-$120k/yr OnsiteFull Time
Visionist
Visionist: Provides software engineering and data analytics for national security.
Active Top Secret/TS-SCI with polygraph required. Bachelor's in a technical discipline (or 4 yrs experience), experience with Python, Kubernetes/Helm, Argo CD, AWS, strong communication and rapid learning ability.
Python, Kubernetes, Helm, Argo CD, AWS
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure
4w
Save
Mark Applied
Hide
Staff AI Infrastructure Engineer
Austin or Reston
HybridFull Time
Seekr
Seekr: Transparent AI platform for enterprise and government decision-making.
8+ YOE8+ years building distributed systems and cloud-native AI infrastructure; strong Python and systems programming skills; Kubernetes, GPU inference, and platform engineering experience; leadership and architecture experience.
Python, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, Infrastructure-as-Code, GitOps, CI/CD, AWS, Azure, Oracle Cloud Infrastructure, Google Cloud Platform
1d
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
1mo
Save
Mark Applied
Hide
Senior AI/ML Engineer
United States or Minneapolis or Washington
$120k-$215k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
3+ YOEBachelor's in CS or quantitative field; 3+ years programming (Python, SQL, packaging), experience with observability, model registry/experiment tracking, responsible AI, GenAI/agents, serving/inference, and cloud platforms.
Python, pytest, SQL, shell, LangChain, Vertex AI, MLflow, AWS, Azure, GPUs, TPUs
1w
Save
Mark Applied
Hide
Model Operations Engineer
Ashburn, Virginia, United States
$108k-$179k/yr HybridFull Time
ManTech
ManTech: Provides technology solutions for defense and intelligence agencies.
Deep ML/Ops experience deploying and monitoring ML models, building training/inference pipelines, working with cloud and container platforms, and collaborating with data science and engineering teams.
MLflow, Kubeflow, Airflow, Alibi, Grafana, AWS SageMaker, Databricks, DataRobot, Docker, AWS, Azure, GCP, Python, Scala, Java, Elastic Stack, Hadoop, Spark, Impala, MySQL, Elasticsearch, Solr, Oracle, Postgres, MongoDB, Lambda, GraphQL, PyTorch, TensorFlow, Keras, OpenCV, SimpleITK, ITK, VTK
1d
Save
Mark Applied
Hide
ML Ops Engineer
Chantilly, Virginia, United States
HybridFull Time
Bana Solutions
Bana Solutions: Provides software engineering, cybersecurity, and data solutions for government agencies.
6+ YOEOwnership of full production ML lifecycle, model serving for low-latency inference, monitoring/drift detection, automated retraining pipelines, Kubernetes/Docker deployment, strong Python and software engineering skills.
MLflow, SageMaker, Weights & Biases, Vertex, PyTorch, TensorFlow, ONNX, Feast, Kubernetes, Docker, Helm, Prometheus, Grafana, Evidently, TorchServe, Triton, KServe, Seldon, BentoML, Apache Airflow, Kafka, Bytewax, Flink, Spark Streaming, Kafka Streams, CUDA, NVIDIA, S3, MinIO, Redshift, EKS, OpenShift, Python
2w
Save
Mark Applied
Hide
AI Integration Engineer
Annapolis Junction, Maryland, United States
$113k-$257k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Consulting and technology services for government and commercial clients
3+ YOE3+ years building production Python systems, experience integrating APIs and LLMs, implementing inference flows and integration tests, knowledge of HTTP auth patterns, containerization and CI/CD, ability to travel up to 25%, TS/SCI with polygraph, bachelor’s in CS/Engineering/Data Science.
Python, HTTP, OAuth2, API keys, JWT, CI/CD, Docker, Kubernetes, Kafka, SQS, SNS, Redis, Postgres, AWS Bedrock AgentCore, Google Gemini Enterprise Agent Platform, Microsoft Foundry Agent Service, LangGraph, CrewAI, AutoGen
1mo
Save
Mark Applied
Hide
AI Engineer (PPA - 009)
Columbia, Maryland, United States
OnsiteFull Time
SageCor Solutions
SageCor Solutions: Provides engineering services for the intelligence community.
15+ YOEActive TS/SCI with polygraph required. 15+ years experience (15 with Master's, 18 with Bachelor's) in CS/Engineering/Data Science/Math; expertise in Python and JavaScript/TypeScript, vector DBs, embeddings, model APIs, model hosting/inference pipelines, secure AI practices, and defense or intelligence support.
Python, JavaScript, TypeScript
1d
Save
Mark Applied
Hide
Senior Software Engineer
Fort Meade, Maryland, United States
$190k-$240k/yr OnsiteFull Time
Belay Technologies
Belay Technologies: Provides specialized technology and engineering services to the DoD.
3+ YOETS/SCI with polygraph required; 3+ years software experience; proficiency with Python, Argo CD, Kubernetes/Helm, and AWS; strong communication, leadership, and mentoring skills; experience with inference/LLM serving is a plus.
Python, Argo CD, Kubernetes, Helm, AWS, vLLM, LiteLLM, Elastic, Grafana, Prometheus, Docker
2w
Save
Mark Applied
Hide
AI/ML Engineer — Generative AI Mission Systems
Laurel or United States
HybridFull Time
Rackner
Rackner: Builds cloud-native software and AI systems for government agencies.
4+ YOE4+ years building LLM, RAG, and prompt-engineered AI; master's or PhD in AI/ML or related; experience with inference pipelines, agentic AI, and delivering LLM-enabled software; strong collaboration and documentation skills.
Kubernetes, OpenShift, WebAssembly (WASM)
3w
Save
Mark Applied
Hide
AI Integration Engineer
Annapolis Junction or McLean
$113k-$257k/yr OnsiteFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Provides technology and management consulting services to diverse organizations.
3+ YOE3+ years building production software in Python, integrating services/APIs and LLMs, implementing inference flows, writing integration tests, and knowledge of HTTP auth patterns and container/CI/CD tooling.
Python, OAuth2, JWT, Kafka, SQS, SNS, AWS Bedrock AgentCore, Google Gemini Enterprise Agent Platform, Microsoft Foundry Agent Service, Docker, Kubernetes, Redis, Postgres, LangGraph, CrewAI, AutoGen
2mo
Save
Mark Applied
Hide
Director, Machine Learning Engineering
Palo Alto or New York City or Dallas or Bethesda or Seattle
$150k-$300k/yr OnsiteFull Time
GEICO
GEICO: Provides vehicle and property insurance services to consumers.
10+ YOE10+ years engineering or AI/ML leadership experience; proven track record building scalable distributed systems and personalization platforms; expertise in RAG, context/memory architectures, and real-time inference; strong business acumen and cross-functional leadership.
Retrieval-Augmented Generation (RAG), LLMs
3w
Save
Mark Applied
Hide
Sr. GenAI Specialist PSA , Solutions Architecture
Herndon, Virginia, United States
$154k-$208k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE7+ years designing and operating distributed applications; 5+ years customer-facing experience building production GenAI/ML systems; hands-on LLM, RAG, agentic workflows, prompt engineering, inference optimization; experience with AWS AI/ML services.
Amazon Bedrock, SageMaker, AgentCore, Bedrock Agents, Strands Agents SDK, LangChain, OpenSearch Serverless, pgvector, Pinecone