1 model deployment engineer job at 1 company in Hampton, VA

PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
Python, PyTorch, vLLM, SGLang, TensorRT, LLMs
1mo
Save
Mark Applied
Hide
Gen AI Engineer
Atlanta or Richmond or Lake Mary or Nashville or Indianapolis or Wilmington or Mason or Grand Prairie or Norfolk or United States
HybridFull Time
Elevance Health
Elevance HealthNYSE: ELV: Provides health insurance plans and integrated healthcare services.
4+ YOEBachelor's in a quantitative field (or equivalent) and 4+ years experience; advanced Python and SQL; hands-on experience with LLM/GenAI, NLP, Hugging Face, TensorFlow, Keras, PyTorch, Spark; cloud (GCP/AWS) and MLOps/model deployment experience.
Python, SQL, Hugging Face, TensorFlow, Keras, PyTorch, Spark, GCP, AWS, vLLM, Text Generation Inference, FastAPI, GPTQ, AWQ, bitsandbytes, AWS EKS, AWS ECS, GCP GKE, Azure AKS