AMD
Posted 3w ago

Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops

AMD
Santa Clara, California, United States
OnsiteFull Time
Responsibilities
  • architecting stack
  • optimizing performance
  • mentoring engineers
Requirements
  • Expert in high-performance C++ and GPU programming (HIP/CUDA)
  • Experience with LLMs and AI systems
  • GPU profiling and kernel optimization, and degree in CS/CE/EE
Technical tools mentioned
C++HIPCUDAROCmAMD ROCm ProfilerNVIDIA NsightCV-CUDAcuDNNNCCLPyTorchTensorFlowJAX

Job description



WHAT YOU DO AT AMD CHANGES EVERYTHING 

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.  Together, we advance your career.  




THE ROLE:

 

AMD is looking for an influential software engineer who is passionate about improving the performance of key applications and benchmarks. You will be a member of a core team of incredibly talented industry specialists and will work with the very latest hardware and software technology.  

 

THE PERSON:

 

As a Staff Software Developer, you will be at the heart of AMD's AI strategy, tackling one of the most exciting challenges in the industry: training and running AI to make AI itself more efficient on GPUs on the fly, which can dramatically alter the trajectory of AI progress. This is a high-impact, hands-on role where your work will directly define the software that powers the future of AI.

 

KEY RESPONSIBILITIES: 

 

Architect and Drive the AI Software Stack: You will establish best practices and optimize performance from the lowest-level GPU kernels to large-scale distributed systems, shaping the foundational software for AMD hardware. By leveraging cutting-edge Large Language Models (LLMs) and agent-based technologies, you will accelerate the development and performance enhancement of the AMD ROCm ecosystem, ensuring it remains at the forefront of AI innovation.


Accelerate Foundational Models: Your work will directly accelerate cutting-edge applications like foundation models (LLMs) and autonomous AI agents, ensuring AMD is the platform of choice for the most demanding workloads.


Innovate Across Hardware and Software: You will contribute to the entire co-design lifecycle, from influencing future GPU architectures to developing groundbreaking software for new accelerators and collaborating with the broader AI community.


Success in this role requires a deep passion for software engineering, strong technical ownership to see complex problems through to resolution, and the ability to influence technical direction across teams. As a senior engineer, you will also be expected to mentor others and effectively communicate your ideas to shape the future of AI at AMD.


To excel in this role, we seek a candidate with exceptional technical expertise, who can bridge deep proficiency in high-performance C++ software engineering and low-level GPU programming with a robust understanding of Large Language Models (LLMs) and AI systems. The ideal candidate can bridge kernel engineering with AI post-training (RL) experience. A great candidate is deep in one and light on the other. 


Kernel engineering means demonstrating mastery in designing complex, scalable systems using modern C++, coupled with a fundamental grasp of GPU architectures (HIP/CUDA), memory hierarchies, and kernel optimization to maximize hardware performance. This expertise should be evidenced by significant hands-on experience in large-scale C++/HIP/CUDA projects, such as contributing to the ROCm ecosystem (e.g., rpp, MIVisionX, rocAL, rocdecode, rocjpeg), CUDA libraries (e.g., CV-CUDA, cuDNN, NCCL), or the C++/HIP/CUDA core of ML frameworks like PyTorch, TensorFlow, or JAX. 


AI post-training is equally critical, and requires deep understanding of LLMs, including but not limited to transformer architectures, attention mechanisms, and the full model lifecycle, with hands-on experience in advanced model alignment and post-training techniques like Supervised Fine-Tuning (SFT) and Reinforcement Learning (e.g., RLHF, GRPO). Candidates must also stay at the forefront of LLM advancements, showing familiarity with cutting-edge trends such as Mixture-of-Experts (MoE) architectures, inference optimizations (e.g., quantization, speculative decoding), and modern application patterns like Agentic AI systems (e.g. AlphaEvolve for code/kernel generation). 


Experience and interest in code generation and/or self-improving LLMs is a plus.

 

PREFERRED EXPERIENCE: 

 

  • This is a senior role that requires a unique blend of expertise across software engineering, GPU computing, and artificial intelligence. The ideal candidate will possess:

    • Lengthy professional software development experience in performance-critical environments.
      Extensive hands-on experience in GPU programming (HIP/CUDA) and optimizing deep learning kernels and operators.
    • Computer vision expertise
    • A fundamental understanding of GPU architecture and memory hierarchy, used to diagnose and resolve complex performance bottlenecks.
    • Expert-level proficiency in modern C++ and object-oriented design.
    • Deep experience using GPU profiling and performance analysis tools (e.g., AMD ROCm Profiler, NVIDIA Nsight) to diagnose and resolve complex bottlenecks in distributed, multi-GPU systems.
    • Deep knowledge of transformer architectures, attention mechanisms, and modern AI systems (Generative AI, Agentic AI).
    • Hands-on experience optimizing the post-training and inference pipelines of Large Language Models (LLMs).
    • Strong technical ownership, communication, and problem-solving skills with a track record of delivering complex technical solutions.
    • Plus: Experience or deep expertise with the AMD ROCm/HIP ecosystem.

 

ACADEMIC CREDENTIALS: 

  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
  • Relevant publications in AI/ML, GPU computing, or system optimization are highly valued.

 

This role is not eligible for visa sponsorship.




Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD’s “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.

About AMD

Designs and manufactures computer processors and graphics technology.

Similar jobs

Software Development Engineer roles near Santa Clara, California
2d
Save
Mark Applied
Hide
Software Dev Engineer, AWS Networking
Santa Clara or Seattle
$165k-$224k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOERequires 3+ years professional software development, 2+ years systems design and architecture, large-scale distributed software experience, and proficiency in C#, C++, Java, Perl, or another programming language.
C#, C++, Java, Perl
3d
Save
Mark Applied
Hide
Sr. Engineer, Cloud Native - AI Detection and Response (AIDR) (Hybrid, Sunnyvale)
Sunnyvale, California, United States
$140k-$215k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years of software development and 5+ years building cloud-native microservices; expertise in Go, Python, or Java; strong distributed systems, cloud architecture, API, data storage, debugging, and AI technology skills.
Go, Python, Java, REST, gRPC, Protocol Buffers, PostgreSQL, Redis, Docker, Kubernetes, AWS, Oracle Cloud Infrastructure (OCI), Google Cloud Platform (GCP), Microsoft Azure, Kafka, Pulsar, Splunk, API Gateway, Golang, Microsoft CoPilot Ecosystem, MS Agent Studio, MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, SHAP, LIME
3d
Save
Mark Applied
Hide
Software Development Engineer
Seattle or San Francisco or San Jose
$139k-$258k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOE5+ years in observability or telemetry; experience with AI, LLMs, web and mobile platforms, JavaScript, TypeScript, Python, Go, CI/CD, cloud and monitoring tools; bachelor's or advanced computer science degree or equivalent.
JavaScript, TypeScript, Python, Go, iOS, Android, CI/CD, Jenkins, Git, Artifactory, Kubernetes, Ethos, Splunk, New Relic, Datadog, OpenTelemetry, Loki, AppDynamics, Prometheus, AWS, Grafana
3d
Save
Mark Applied
Hide
Software Development Engineer
San Francisco or Denver
$155k-$186k/yr HybridFull Time
Fastly
FastlyNYSE: FSLY: Provides edge cloud platform for content delivery and cybersecurity.
1+ YOERequires 1–3 years of development experience, GoLang knowledge, cloud and containerization exposure, infrastructure-as-code experience, distributed systems understanding, and strong communication skills.
Terraform, Jenkins, Kubernetes, Chef, GoLang, AWS, GCP, Docker, Envoy, API Gateways, Varnish, VCL
1w
Save
Mark Applied
Hide
Software Development Engineer Graduate (Intent-Based Networking) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a related technical discipline; networking knowledge; front-end technologies; one back-end language; SQL and NoSQL database skills; strong coding and design practices.
Software Defined Networking (SDN), HTML, CSS, JavaScript, TypeScript, Python, Go, C, C++, Rust, SQL, NoSQL, Robotron, Apstra, Forward Networks, TCP/IP, DNS, ARP
1w
Save
Mark Applied
Hide
Software Development Engineer Graduate (TikTok - Testing - Growth) - 2027 Start
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
$128k-$256k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Bachelor's or master's degree in a relevant computing field, programming skills in Java, Python, Objective-C, or Golang, testing fundamentals, and knowledge of algorithms and system architecture.
Java, Python, Objective-C, Golang, API, Docker, K8s, Kafka, Redis, Django, Flask
1w
Save
Mark Applied
Hide
Software Development Engineer - Equipment Control
Sunnyvale, California, United States
$160k-$271k/yr OnsiteFull Time
Intuitive
IntuitiveNASDAQ: ISRG: Robotic-assisted systems for minimally invasive surgery.
3+ YOEBachelor's degree plus 5 years or master's plus 3 years in engineering, with software design and development experience. Requires programming, software architecture, hardware integration, systems design, and technical leadership skills.
C#, Python, Programmable Logic Controllers (PLCs), Modbus, OPC, EtherNet/IP, Software Development Life Cycle (SDLC)
1w
Save
Mark Applied
Hide
Sr Software Development Engineer (Full Stack) - Evisort AI
Pleasanton, California, United States
$176k-$264k/yr HybridFull Time
Workday
WorkdayNASDAQ: WDAY: Provides cloud-based software for financial and human capital management.
8+ YOERequires 8+ years of software engineering, 5+ years with web frameworks and relational databases, plus cloud SaaS, distributed systems, DevOps, containers, and cross-functional collaboration experience.
Python, Flask, Django, FastAPI, PostgreSQL, AWS, DevOps, CI/CD, Docker, Kubernetes, Elasticsearch, Java, TypeScript