Together AI
Posted 2w ago

Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking

Together AI
Singapore or San Francisco
HybridFull Time
Responsibilities
  • optimizing inference
  • fine-tuning
  • supporting customers
Requirements
  • 5+ years experience with inference systems
  • Open-source LLM deployment, and post-training pipelines
  • Expert with inference engines
  • Strong Python skills
  • Mandarin and English proficiency
Technical tools mentioned
vLLMTensorRT-LLMSGLangPythonLoRASFTDPORLHFGRPOFlashAttentionHyenaFlexGenRedPajama

Job description

About the role

As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to our most strategic customers — production AI teams looking to leverage high quality models and do inference at scale. For us, FDE is not a replacement for a Solutions Architect; you will partner with our SAs as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment. As key contributors to both the CX, Engineering, and Sales organizations, FDEs add tremendous value by ensuring we can meet the requirements of our most complex POCs, facilitate successful platform adoption, and guide tailored optimization efforts — directly impacting customer success, company growth, and the hardening of our core platform.

Must be a permanent resident or citizen of Singapore. 

Responsibilities

  • Inference Engine Optimization: Select, configure, and optimize inference engine based on hardware, model architecture, and workload profile
  • Configuration & Performance Tuning: Develop configuration updates to win critical POCs, benchmarks, and optimize customer deployments; tune KV cache, apply speculative decoding, determine optimal tensor parallelism, and determine quantization strategy to hit throughput and latency targets.
  • Post-Training & Fine-Tuning: Drive hands-on RL training runs and optimize system design; guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
  • Strategic Customer Alignment: Act as the primary technical point of contact for aligned strategic accounts — monitoring and optimizing endpoint configurations, helping customers get the most out of the platform, and collaborating to ensure we hit critical milestones.
  • Opinionated Onboarding: Establish direct alignment with strategic customers at onboarding; ensure the right inference and post-training configurations are in place from day one to improve time-to-value.
  • Product Feedback Loop: Directly influence our software and model roadmap by surfacing insights from the field. Contribute back to the product where needed to support customer requirements or drive a better experience. Drive early feature and research adoption with strategic logos.

Qualifications

  • Experience: 5+ years in a technical role, with a strong focus on inference systems, open-source LLM deployment, or post-training workflows.
  • Inference Engine Depth: Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang); ability to diagnose and resolve performance issues at the engine level.
  • Inference Optimization: Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques
  • Post-Training Knowledge: Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO; ability to advise on system design
  • Model Landscape Awareness: Broad knowledge of state-of-the-art open-source models and strong judgment on model selection for specific customer use cases, hardware profiles, and performance targets.
  • Coding Proficiency: Strong Python skills; comfortable working in production environments

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers on our journey in building the next generation of AI infrastructure. 

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. 

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our Privacy Policy at https://www.together.ai/privacy

About Together AI

Cloud platform for training and deploying artificial intelligence models.

Year founded
2022
Employees
335
Organization type
Private
Latest investment
Raised $305.00M Series B (2025) — led by General Catalyst, Prosperity7 Ventures
Headquarters
US

Similar jobs

Forward Deployed Engineer roles near San Francisco, California
5h
Save
Mark Applied
Hide
Fouding Forward Deployed Engineer - US
Austin or New York City or San Francisco or Paris or London
HybridFull Time
H
H: Develops autonomous AI agents and foundational action-oriented models.
3+ YOERequires 3+ years of AI and software engineering experience, Python fluency, AI/ML and LLM integration, TypeScript, backend/frontend frameworks, production systems, API/database integration, communication, and customer-facing skills.
Python, TypeScript, FastAPI, Next.js, React
1d
Save
Mark Applied
Hide
Forward Deployed Engineer, Infrastructure Specialist (France)
France or Europe or Munich or Paris or Toronto or London or New York City or San Francisco or Montreal or Berlin or Seoul
HybridFull Time
Cohere
Cohere: Provides enterprise-grade large language models and AI software platforms.
Customer-facing enterprise software deployment experience in private or hybrid clouds; Kubernetes, Helm, DevOps, CI/CD, cloud infrastructure, networking, virtualization, and fluent French required.
Kubernetes, Helm, Git, Azure, AWS, GCP, CI/CD
1d
Save
Mark Applied
Hide
Forward Deployed Engineer (Training)
San Francisco, California, United States
$200k-$400k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
1+ YOERequires 1–2 years of software engineering experience, production debugging, ambiguous-problem ownership, customer communication, and interest in AI inference and training. On-call availability required.
Kubernetes, Slurm, Ray, vLLM, TensorRT-LLM, SGLang, PyTorch, JAX, InfiniBand, RoCE
2d
Save
Mark Applied
Hide
Forward Deployed Engineer
Seattle or San Francisco
OnsiteFull Time
SageOx
SageOx: Shared context infrastructure for AI-native teams.
Software engineering experience with end-to-end integration debugging and independently shipping features. Requires customer-facing communication, clear writing, AI tooling familiarity, and willingness to travel.
Slack
2d
Save
Mark Applied
Hide
Forward Deployed Engineering, Portworx
Santa Clara, California, United States
$208k-$305k/yr OnsiteFull Time
Pure Storage
Pure StorageNYSE: PSTG: Provides all-flash enterprise data storage and management solutions.
5+ YOERequires 5+ years in infrastructure, platform engineering, SRE, or field/solutions engineering; deep production Kubernetes expertise; enterprise storage, Linux, on-prem/cloud, virtualization, technical writing, and relationship-building skills.
Kubernetes, CaaS, Portworx, KVDB, VMware, KubeVirt, SAN, NAS, Fibre Channel, iSCSI, NVMe/TCP, dm-multipath, LVM, Linux, CI/CD, CSI
2d
Save
Mark Applied
Hide
Forward Deployed Engineer
United States or Palo Alto
$150k-$170k/yr HybridFull Time
Spotnana
Spotnana: Cloud-based travel-as-a-service platform for enterprises.
3+ YOERequires 3+ years of software engineering experience, distributed systems, REST APIs, webhooks, cloud infrastructure, debugging, observability, and direct enterprise customer or partner experience.
REST APIs, AWS, GCP, Azure, Datadog, PagerDuty, RocketLawyer, International Airlines Travel Agent Network (IATAN), Microsoft Teams, LinkedIn
2d
Save
Mark Applied
Hide
Member of Technical Staff - Forward Deployed Engineer
San Francisco, California, United States
$200k-$275k/yr OnsiteFull Time
Halluminate
Halluminate: Building reinforcement learning environments for training AI agents.
2+ YOE2+ years in consulting or forward-deployed engineering for large enterprises; strong engineering and customer communication skills; 0-to-1 delivery experience; ability to translate customer needs into technical requirements; AI agent evaluation experience preferred.
Excel
3d
Save
Mark Applied
Hide
Forward Depolyed Engineer - RiskOS Agents
United States or San Francisco or Seattle or New York City or Carson City
$250k-$280k/yr HybridFull Time
Socure
Socure: Provide AI-driven identity verification and fraud prevention software.
4+ YOE4+ years in hands-on software, forward deployed, solutions, implementation, applied AI engineering, technical consulting, or similar roles; experience with APIs, integrations, production systems, and customer-facing problem solving.
Python, SQL, APIs, LLM workflows, retrieval systems, evaluation frameworks, observability