📋 External Recruiting Agencies

Bright Vision Technologies is a staffing and IT consulting firm that recruits talent for clients and provides H-1B sponsorship services, as evidenced by its own self-description and job postings mentioning direct clients.

This company was flagged and excluded from default search results. Proceed with caution.

Bright Vision Technologies
Posted 6d ago

Machine Learning Infrastructure Engineer

Bright Vision Technologies
United States
$105k-$143k/yrRemoteFull Time
Responsibilities
  • designing platforms
  • operating services
  • optimizing inference
Requirements
  • Bachelor's or master's in computer science or related field
  • 6+ years in distributed systems
  • Infrastructure, or ML platforms
  • Python and Go, Rust, or C++
  • Production-scale inference
  • Kubernetes, cloud, GPUs, and observability
Technical tools mentioned
PythonGoRustC++vLLMTensorRT-LLMKubernetes

Job description

Machine Learning Infrastructure Engineer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Machine Learning Infrastructure Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $105,000–$143,000 Annually
Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary
We are seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving.

Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.
Preferred Qualifications
  • Open-source contributions to model serving infrastructure.
  • Experience with multi-region or globally distributed AI serving.
  • Familiarity with model quantization, distillation, and compression techniques.
  • Exposure to FinOps for AI workloads and cost-efficient serving design.
  • Experience supporting external-facing AI APIs at scale.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected].
Bright Vision Technologies is an Equal Opportunity Employer.

Similar jobs

Machine Learning Infrastructure Engineer roles
5d
Save
Mark Applied
Hide
Senior Machine Learning Infrastructure Engineer, Embedding Platform
United States
$191k-$267k/yr RemoteFull Time
Reddit
RedditNYSE: RDDT: Social platform for community-driven discussion and content sharing.
5+ YOERequires 5+ years in machine learning engineering, deep learning and distributed ML expertise, Python and modern ML frameworks, production-scale infrastructure, system design, testing, optimization, evaluation, and strong communication.
Python, PyTorch, TensorFlow
6d
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer, Technology
New York City, New York, United States
$185k-$300k/yr OnsiteFull Time
Point72
Point72: Global alternative investment firm managing capital and venture investments.
3+ YOEBachelor’s or master’s degree in a technical field; 3–7 years building scalable compute or ML infrastructure; expertise in distributed systems, Kubernetes, cloud platforms, MLOps, Python, and systems programming.
Kubernetes, AWS, Google Cloud Platform, Azure, MLflow, Ray, Airflow, Kubeflow, Terraform, Python, Go, C++, Rust
6d
Save
Mark Applied
Hide
Tech Lead, Machine Learning Infrastructure Engineer
Seattle or Bellevue
$198k-$416k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
5+ YOEBachelor’s or master’s degree in a technical discipline; 5+ years software engineering and 3+ years ML infrastructure experience. Requires Python, C++ or Java, distributed systems, GPU clusters, Kubernetes or Slurm, PyTorch or TensorFlow, CUDA, and networking expertise.
Python, C++, Java, Kubernetes, Slurm, PyTorch, TensorFlow, CUDA, InfiniBand, RoCE, Megatron, DeepSpeed, Triton, Cutlass, TensorRT, Triton Inference Server
1w
Save
Mark Applied
Hide
Staff ML Infra Engineer
Mountain View, California, United States
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides an end-to-end e-commerce and logistics network.
5+ YOEBachelor's in CS/EE/math,5+ years applied ML experience,proficiency in Python/Java,experience with big data pipelines,ML frameworks,cloud platforms,and building scalable low-latency services.
Python, Java, Hadoop, Hive, Presto, Spark, Scala, Apache Airflow, TensorFlow, PyTorch, AWS, GCP
3w
Save
Mark Applied
Hide
Staff Engineer - ML Infra / MLOps
Palo Alto, California, United States
$218k-$285k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
3w
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer, Safeguards Research
San Francisco or New York City
$350k-$500k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Proven software engineering with Python, experience building and operating data-intensive or distributed systems, tooling for researcher workflows, debugging performance and correctness, and strong communication skills.
Python, transformers, GPU
4w
Save
Mark Applied
Hide
Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)
Toronto or Zürich or California
HybridFull Time
Veeda AI
Veeda AI: Building multimodal foundation world models for physical AI.
Bachelor's degree or equivalent experience; deep experience with distributed deep-learning training, numerical precision (BF16/FP8), hardware/software root-cause analysis, and strong Python and C++/CUDA skills.
PyTorch, FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism, BF16, FP8, Python, C++/CUDA, NCCL, ROCm, JAX/XLA, CUDA, Triton
1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker