Graphcore
Posted 8mo ago

Senior Machine Learning Engineer (Large Systems)

Graphcore
Cambridge or Bristol or Gdańsk or London
OnsiteFull Time
Responsibilities
  • implementing models
  • optimising performance
  • benchmarking models
Requirements
  • Advanced ML engineering skills with deep learning experience
  • Distributed training across many accelerators
  • Strong Python or C++ development
  • Proficiency with PyTorch/JAX, and ability to benchmark and optimise models for performance
Technical tools mentioned
PyTorchJAXPythonC++TritonCUDAKubernetesInfinibandNVLinkRoCE

Job description

About Graphcore

At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.

Job Summary

As a Senior Machine Learning Engineer in the Applied AI team at Graphcore, you will contribute to advancing AI technology by developing and optimising AI models tailored to our specialised hardware. You will work on large scale systems where performance is critical to the success of our projects. Working closely with the Software development and Research teams, you will play a critical role in identifying opportunities to innovate and differentiate Graphcore’s technology. We seek engineers with strong technical skills and an understanding of AI model implementation at scale, eager to make a tangible impact in this rapidly evolving field.


The Team

The Applied AI team’s role is to be proxies for our customers, we need to understand the latest AI models, applications, and software to ensure that Graphcore’s technology works seamlessly with the AI ecosystem and at scale. We build reference applications, contribute to key software libraries e.g. optimising kernels for efficiency on our hardware, and collaborate with the Research team to develop and publish novel ideas in domains such as efficient compute, model scaling and distributed training and inference of AI models for multiple modalities and applications.

If you're excited about advancing the next generation of AI models on cutting-edge hardware, we’d love to hear from you!

Responsibilities and Duties

  • Implement latest machine learning models and optimise them for performance and accuracy, scaling to 1000s of accelerators.
  • Test and evaluate new internal software releases, provide feedback to software engineering teams, make necessary code fixes, and conduct code reviews.
  • Benchmark models and key ML techniques to identify performance bottlenecks and improve model efficiency.
  • Design and conduct experiments on novel AI methods, implement them and evaluate results.
  • Collaborate with Research, Software, and Product teams to define, build, and test Graphcore’s next generation of AI hardware.
  • Engage with AI community and keep in touch with the latest developments in AI.

  

Candidate Profile

Essential:

  • Bachelor/Master's/PhD or equivalent experience in Machine Learning, Computer Science, Maths, Data Science, or related field.
  • Proficiency in deep learning frameworks like PyTorch/JAX.
  • Strong Python or C++ software development skills
  • Expertise in deep learning from model training to optimisation and evaluation.
  • Experience in distributed training or inference of ML models across 64+ accelerators.
  • Capable of designing, executing and reporting from ML experiments.
  • Developed deep understanding of performance bottlenecks and how to overcome them.
  • Ability to move quickly in a dynamic environment
  • Enjoy cross-functional work collaborating with other teams.
  • Strong communicator - able to explain complex technical concepts to different audiences.

Desirable:

  • Experience in one or more of:
    • MLOps for Kubernetes-based clusters
    • Building production systems with large language models
    • Efficient computing based on low-precision arithmetic.
  • Experience writing C++/Triton/CUDA kernels for performance optimisation of ML models.
  • Familiarity with HPC systems and networking including Infiniband, NVLink, RoCE technologies.
  • Have contributed to open-source projects or published research papers in relevant fields.
  • Knowledge of cloud computing platforms.
  • Keen to present, publish and deliver talks in the AI community.

Benefits

In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.

Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications

About Graphcore

Produces specialized processors for artificial intelligence workloads.

Year founded
2016
Employees
450
Organization type
Private
Latest investment
Raised $222.00M Series E (2020) — led by Ontario Teachers' Pension Plan, Fidelity International, Schroders
Headquarters
GB

Similar jobs

Machine Learning Engineer roles near Cambridge, England
6h
Save
Mark Applied
Hide
2027 - Full Time Assistant Vice President - Corporate Functions, Analytics Lab (Machine Learning Engineer)
Jersey City or Montreal or Paris or Lisbon or London or New York or Chesterbrook or San Francisco or Boston or Chicago or Denver or Miami or Washington
$150k/yr OnsiteFull Time
BNP Paribas
BNP ParibasEuronext Paris: BNP: Provides global banking, financial, and investment services.
Master’s degree in a quantitative field; December 2026–June 2027 graduation; Python or C, SQL, statistical modeling, data analysis, LLM/agentic workflows, ML frameworks, RAG systems, Git, and communication skills.
Python, C, SQL, LangChain, CrewAI, Google ADK, OpenAI Agent SDK, TensorFlow, PyTorch, scikit-learn, Git, DVC, MLflow, Jenkins, Docker
1d
Save
Mark Applied
Hide
Staff Machine Learning Engineer
London or United Kingdom
HybridFull Time
BJAK
BJAK: Online platform for insurance comparison and road tax renewal.
Experience shipping real machine learning systems, working with large models, writing production-grade code, and owning scalable, reliable ML systems and outcomes.
Python, PyTorch, JAX, LoRA, QLoRA, SFT, DPO
2d
Save
Mark Applied
Hide
Senior Machine Learning Engineer
London, England, United Kingdom
OnsiteFull Time
AECOM
AECOMNYSE: ACM: Provides infrastructure consulting and engineering services for global projects.
Requires production ML experience, a master's degree or PhD, Python and modern ML frameworks, model design and optimization expertise, independent delivery, and strong stakeholder communication.
Python, PyTorch, TensorFlow, scikit-learn, MLOps, SaaS
3d
Save
Mark Applied
Hide
Senior Machine Learning Engineer, TikTok BRIC Community Health
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
$162k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOEMaster's degree or above in a relevant technical field and 2+ years of machine learning experience. Requires strong software engineering, machine learning, Python or Java/C++/Go, and Spark, Hadoop, or Hive experience.
Python, Java, C++, Go, Spark, Hadoop, Hive, LLMs
4d
Save
Mark Applied
Hide
Member of Technical Staff (Machine Learning Engineer, Ranking Quality - Search)
Belgrade or London or Berlin
OnsiteFull Time
Perplexity
Perplexity: AI-powered search engine providing conversational answers with citations.
5+ YOE5+ years of relevant industry experience; deep search or recommender-system expertise; production ranking ownership; strong machine-learning and software-engineering skills; depth in neural ranking or low-latency ranking systems.
4d
Save
Mark Applied
Hide
Senior Machine Learning Engineer - FinCrime
London, England, United Kingdom
£88k-£111k/yr HybridFull Time
Wise
WiseLondon Stock Exchange: WISE: Online platform for sending and managing money internationally.
STEM degree, strong statistics, ML lifecycle experience, Python or Java, advanced SQL, data pipelines, analysis, visualization, and model deployment experience.
Python, Java, SQL, Kaggle, KDD, Google Summer of Code (GSoC), Graph Neural Networks (GNNs), Support Vector Machines (SVM), Natural Language Processing (NLP), Transformers, LSTMs, Kafka
5d
Save
Mark Applied
Hide
Senior Machine Learning Engineer
London, London, United Kingdom
RemoteFull Time
Ravelin
Ravelin: Machine learning platform for online payment fraud detection.
Production ML systems experience, scalable training pipelines, PyTorch or TensorFlow, distributed multi-GPU training, transformer architectures, workflow orchestration, software engineering, and cross-functional technical leadership.
PyTorch, TensorFlow, Prefect, Kubeflow, Argo, Git, CI/CD, Go, C++, Java, Rust, dbt, Transformers
5d
Save
Mark Applied
Hide
Machine Learning Engineer, Memory
London or Cary
OnsiteFull Time
Epic Games
Epic Games: Develops video games and licenses the Unreal game engine.
3+ YOEPhD in computer science, mathematics, or related field, or 3+ years relevant industry experience; machine learning production systems experience; expertise in foundation models, memory, Python, and ML frameworks.
Python, NumPy, SciPy, scikit-learn, PyTorch