📋 External Recruiting Agencies

Jobgether is a remote job matching platform and talent marketplace; the job listing explicitly states it is posted on behalf of an anonymous partner company, making Jobgether a recruiting intermediary rather than the direct employer.

This company was flagged and excluded from default search results. Proceed with caution.

J
Posted 6d ago

AI Optimization Engineer

Jobgether
United States
$100k/yrRemoteFull Time
Responsibilities
  • optimizing workloads
  • profiling systems
  • mentoring engineers
Requirements
  • Bachelor's degree and 6+ years in performance engineering or ML systems. Requires Python, C++
  • GPU optimization
  • Distributed training/inference
  • Profiling
  • Systems optimization, and U.S. work authorization
Technical tools mentioned
PythonC++vLLMTensorRT-LLMDeepSpeedTritonCUTLASSFinOps

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States.

This is a fully remote opportunity focused on improving the performance, scalability, and economics of large-scale AI systems.
You will optimize training and inference workloads across the stack, from low-level GPU kernels to distributed infrastructure.
The role combines systems engineering, performance analysis, machine learning infrastructure, and compiler-level optimization.
You’ll work with modern GPUs and large neural networks, using rigorous measurement and profiling to identify and resolve performance bottlenecks.
The position offers the opportunity to influence production AI workloads where improvements in throughput, latency, and cost have meaningful business impact.
You’ll collaborate closely with engineering, product, operations, and business teams while contributing to technical direction and engineering standards.
As a senior technical contributor, you’ll also mentor engineers and help drive a culture of measurable, production-ready optimization.



Accountabilities:
  • Optimize training and inference workloads to maximize throughput, minimize latency, and improve cost efficiency across large-scale neural network systems.
  • Analyze and improve performance across the full technology stack, including GPU kernels, memory management, communication, distributed systems, and model execution.
  • Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis to guide optimization decisions.
  • Design and implement performance improvements using Python, C++, and relevant AI systems technologies.
  • Optimize distributed training and inference architectures, including model parallelism, communication strategies, and resource utilization.
  • Evaluate and implement model compression techniques while carefully considering their impact on model accuracy and production performance.
  • Investigate complex performance and reliability issues through systematic debugging, instrumentation, benchmarking, and root-cause analysis.
  • Contribute to production-scale optimization of large language model inference and other demanding AI workloads.
  • Develop and improve low-level optimization techniques, including custom GPU kernels where appropriate.
  • Collaborate with product, design, engineering, operations, and business stakeholders to translate ambiguous requirements into scalable, well-engineered technical solutions.
  • Participate in architecture and code reviews, establish engineering best practices, and contribute to long-term technical strategy.
  • Mentor junior and mid-level engineers, helping raise technical quality and strengthen performance engineering capabilities.
  • Identify opportunities to improve the cost structure of AI workloads through infrastructure optimization and FinOps-oriented analysis.
  • Requirements:

    • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline.
    • 6+ years of professional experience in performance engineering, machine learning systems, high-performance computing, or a closely related field.
    • Strong programming proficiency in Python and C++, with the ability to develop production-quality, maintainable code.
    • Hands-on experience optimizing deep learning workloads on modern GPU architectures.
    • Deep understanding of distributed training and inference techniques, including parallelism strategies and communication primitives.
    • Strong knowledge of memory hierarchies, GPU/CPU performance characteristics, and systems-level optimization.
    • Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.
    • Familiarity with model compression methods and their implications for accuracy, performance, and production deployment.
    • Excellent measurement, debugging, analytical reasoning, and problem-solving abilities.
    • Strong communication and collaboration skills, with the ability to explain complex technical concepts to cross-functional stakeholders.
    • Demonstrated ability to work independently, make data-driven technical decisions, and deliver meaningful improvements in production environments.
    • Experience with production-scale LLM inference is strongly preferred.
    • Contributions to projects such as vLLM, TensorRT-LLM, DeepSpeed, or comparable AI systems projects are a plus.
    • Experience with custom kernel development using technologies such as Triton or CUTLASS is preferred.
    • Familiarity with FinOps and cost optimization for AI workloads is advantageous.
    • Publications, conference presentations, or technical talks focused on AI systems or performance engineering are a plus.
    • Must be currently based in the United States and authorized to work in the U.S.; U.S. citizens, permanent residents, EAD holders, and candidates eligible for H-1B transfer are encouraged to apply. New H-1B sponsorship is not available.
    • Benefits:

      • $100,000 annual salary for this full-time direct W2 position.
      • 100% remote work within the United States.
      • Opportunity to work on challenging AI optimization and high-performance computing problems.
      • Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and production AI infrastructure.
      • Significant opportunities for technical ownership, mentorship, and career growth.
      • Collaborative environment spanning engineering, product, operations, design, and business teams.
      • Opportunity to contribute to impactful production AI systems and advance performance, scalability, and cost efficiency.
      • Equal employment opportunity and an inclusive workplace committed to fair treatment of employees and applicants.


How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 Why Apply Through Jobgether? 
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
#LI-CL1

Similar jobs

AI Optimization Engineer roles
2mo
Save
Mark Applied
Hide
AI Optimization Engineer
Houston, Texas, United States
OnsiteFull Time
Occidental Petroleum
Occidental PetroleumNYSE: OXY: Explores for and produces oil and gas resources.
Master's in engineering/applied mathematics, strong optimization/control and ML knowledge, Python and Git proficiency, experience with system modeling and advanced control (MPC), strong communication and independent work skills.
Python, Git, AWS, Docker, Kubernetes
4w
Save
Mark Applied
Hide
AI Flight Optimization Engineering Intern
Oklahoma City, Oklahoma, United States
OnsitePart Time, Internship
Skydweller Aero
Skydweller Aero: Developing solar-powered autonomous aircraft for perpetual uncrewed flight.
Currently enrolled undergraduate/graduate student in CS, aerospace, or related; U.S. Person required; Python and Git experience; familiarity with AI/ML concepts; prior R&D experience preferred.
Python, Git
3mo
Save
Mark Applied
Hide
Applied AI & Optimization Engineer
San Diego or United States
$140-$185/yr HybridFull Time
Firestorm Labs
Firestorm Labs: Develops and manufactures low-cost 3D-printed unmanned aerial vehicles.
5+ YOE5+ years in engineering with optimization/ML systems; Python; MILP/OR-Tools; productionizing algorithms; domain collaboration.
Python, Optimization frameworks, MILP solvers, Constraint solvers, OR-Tools, LLMs
2y
Save
Mark Applied
Hide
AI Performance Optimization Engineer
New York City or San Francisco or London
$120k-$240k/yr HybridFull Time
Lightning AI
Lightning AI: Unified platform to build, train, and deploy AI models.
Expert in deep learning, compiler internals, CUDA/Triton, distributed systems; open-source contributions; degree in CS/Engineering; strong collaboration skills.
PyTorch, JAX, TensorFlow, CUDA, Triton, GPU programming
3mo
Save
Mark Applied
Hide
Edge AI/Model Optimization Engineer
Aberdeen, Maryland, United States
HybridFull Time
NextGen Federal Systems
NextGen Federal Systems: Provides advanced software engineering and mission support to federal agencies.
5+ YOEExperience deploying ML on edge GPUs, optimizing LLMs, and performance tuning; requires security clearance.
CUDA, TensorRT, ONNX Runtime, vLLM, Ollama, Linux, Docker, Kubernetes, Python