Roboflow
Posted 1mo ago

Machine Learning Engineer - Inference Maintainer & Developer Experience

Roboflow
New York City or San Francisco or United States or Europe
RemoteFull Time
Responsibilities
  • maintaining inference
  • building pipeline
  • designing tests
Requirements
  • 5+ years building and operating production ML systems
  • Strong CV/inference foundation
  • CI/CD and test infra experience
  • Proficiency with PyTorch/TensorFlow/ONNX/TensorRT/vLLM, and experience with image/video processing tools
Technical tools mentioned
inferencePyTorchTensorFlowONNXTensorRTvLLMOpenCVDeepStreamPillowPyAVGitHubCI/CD

Job description

Our mission is to make the world programmable. Sight is one of the key ways we understand the world, and soon this will be true for the software we use, too.

We’re building the tools, community, and resources needed to make the world programmable with artificial intelligence. Roboflow simplifies building and using computer vision models. Today, over 1M+ developers, including those from half the Fortune 100, use Roboflow’s machine learning open source and hosted tools. That includes counting cells to accelerate cancer research, improving construction site safety, digitizing floor plans, preserving coral reef populations, guiding drone flight, and much more.

Our team is small relative to our impact, and we believe our user success is our success (not the inverse). A team member summarized: “Roboflow is a company full of giant brains and tiny egos.” We find software has a multiplier effect on all roles (not only product and engineering), so Roboflow employs developers across the company in design, sales, customer support, marketing, and beyond.

We’re supported by great customers and investors, having raised over 63 million from Google Ventures, Y Combinator, Craft Ventures, Sam Altman, Lachy Groom, amongst other leading software investors.

At the center of all of this is inference — one of our most important open source projects and the engine that runs computer vision models everywhere, from cloud GPUs to edge devices in the field. It powers our commercial platform and is relied on by tens of thousands of developers. This role exists to be its steward.

Why This Role Exists

Inference is growing fast — and so is the volume of contributions, increasingly authored with the help of AI agents. That's a great problem to have, but it's outpacing our ability to keep quality high and cut releases on a predictable cadence. Today we ship roughly weekly, and it's a fight.

We want to flip that equation. The goal is to build and continuously evolve an agentic-driven contribution and release pipeline — automated and semi-automated review, triage, CI/CD, and end-to-end testing — so that we can safely absorb a high volume of agent-generated PRs while staying firmly in control of quality. The ideal end state: nightly end-to-end tests across every target (both standalone and on-platform), backed by a growing, world-grounded suite that validates the real health of every build. With that foundation, daily releases become routine, and we can say "yes" to far more contributions without ever lowering the bar — pushing back, by design, according to strictly defined review standards.

Alongside that, this person becomes the human face of inference: teaching internal teams and customers how to get more out of it, partnering with marketing to tell its story, and owning the (genuinely fun) work of bringing new models into the engine.


What We're Looking For

Primarily, you like to make great things with passionate colleagues. You are someone who likes to own outcomes, not only inputs. You're motivated by having responsibility and accountability. You're eager to 'do the work,' big and small.

You're motivated by the question, "How can I improve this?" and have a track record of doing so, even in ways adjacent to your role. Much of our current team is made up of former founders who thrive in the level of autonomy at Roboflow. Maybe you had a side hustle in high school or college.

You care about open source and the developers who depend on it. One of the best ways to stand out among other applicants is to write about something you've built with Roboflow, or to contribute to one of our open source projects — inference especially.

What You'll Do

  • Build and maintain inference, our flagship open source and commercial CV inference engine, keeping it healthy and high-quality as contribution volume scales.

  • Build an agentic-driven contribution pipeline — automated and semi-automated review, triage, and CI/CD — so we can safely accept a high volume of agent-generated PRs and move from weekly releases toward daily ones.

  • Design and grow a world-grounded, ever-expanding test suite that validates real build health across every target (standalone and on-platform), with the goal of nightly end-to-end runs across all of them.

  • Define and enforce the "rules of the road" — the review standards and skills that agents and contributors must follow. Exercise sharp judgment on when to merge fast and when to push back, and encode that judgment into the system itself.

  • Streamline how new models get added to inference (the most fun part of the job) — making it dramatically faster and easier to bring the latest computer vision and ML models to our users.

  • Teach and enable internal teams and customers. Keep our Field Engineers and Support team a step ahead so they can self-serve and go deeper, and help customers get the full value of the product.

  • Be the bridge between core engineering and clients — translating new capabilities into docs, demos, stories, and launches which would help people use inference more effectively.

  • Contribute to and grow the broader open source community around the project.

Who You Are

You are an experienced Machine Learning practitioner who wants to be an important part of an exceptional team that focuses on using Roboflow's computer vision tools to impact and improve every industry. You have high agency and a bias toward action.

  • 5+ years of hands-on experience building and operating production-grade ML systems, ideally involving large-scale deployment of modern AI models.

  • A real CV/ML foundation — you understand what inference does: how computer vision models work internally, how they're deployed across diverse environments, and how to adapt them for real-world, high-impact use.

  • Stellar agentic skills. You build with AI coding agents fluently and have a track record of using them not just to ship features, but to automate the engineering process itself — review, triage, testing, and CI. You have strong instincts for where agents excel and where they need guardrails.

  • Strong CS and systems background, with the ability to independently tackle complex programming, architecture, and reliability challenges and exercise sound judgment on when to move fast and when rigor is essential.

  • Hands-on experience with CI/CD, release engineering, and test infrastructure — you've built or substantially improved automated testing and delivery pipelines before.

  • Practical expertise with core ML technologies, including several of the following: PyTorch, TensorFlow, ONNX, TensorRT, vLLM (or other LLM/model deployment tools).

  • Strong proficiency in image and video processing, including several of the following: OpenCV, DeepStream, Pillow, PyAV, hardware-accelerated video decoding. Experience with video streaming protocols is an advantage.

  • Excellent communication and soft skills. You can teach, write clearly, and collaborate across engineering, support, field, and marketing — and you actually enjoy it. You're comfortable being a public-facing voice for a project.

  • Open source maintenance experience is a strong plus — you know what it takes to steward a busy repo and a community of contributors.

  • Level-up your performance with AI agents.

Where You'll Work

Roboflow is distributed across the US and Europe. We currently have Hubs in New York City and San Francisco (and plan to open more as we grow density in new cities). We provide opportunities (like team onsites in different cities) and resources (like a $4000/yr travel stipend) to work in person with other team members as much as you'd like, while also supporting remote team members. You can work from one of our Hubs (we offer a relocation bonus), work from home, work at co-working spaces, etc. We want you to work where you work best!

What You'll Receive

To determine your salary, we use a number of market and data-driven salary sources. We review all salaries every six months to ensure we stay in line with the market.

💰 We use Tier 1 rates for employees who work out of our San Francisco & New York hubs more than 3+ times per week.

📈 In addition to our cash compensation, we offer generous perks and benefits. Below are some of the highlights:

  • $4000/yr Travel Stipend to travel anywhere anytime to work alongside other Roboflowers

  • $350/mo Productivity stipend to spend on things that make your work environment more productive, like high-speed internet at home or a co-working space

  • $350/mo AI Tools stipend

  • Cover up to 100% of your health insurance costs for you and your partner or family

  • $150/mo team lunch stipend

  • Remote first/flexible schedule allowing you to work collaboratively with other team members and asynchronously

  • Unlimited PTO- with an annual 2 week minimum, we encourage you to take time off for yourself

  • 12 weeks parental leave

  • Equity in the company so we are all invested in the future of computer vision

Interview Process (~5 hours)

Below is the interview process you can expect for this role.

Before the Interview:

  • We’ll review your application, LinkedIn, Github, etc.

  • The best way to stand out is to write about something you’ve built with Roboflow or contribute to one of our open source projects.

  • We may send you a technical screen if applicable.

Introduction Phase:

  • [15m] Technical Assessment

Team Interview Phase:

  • Live coding [45m]

  • Home assignment

  • [30m] Meet with Inference Core team member

  • [60m] Meet with hiring manager

    • Use this time to review specifics about the job description

    • Begin working through your 30/60/90 projects

    • Ask questions!

Final Interview Stage:

  • [45m] Meet with Head of Operations for a culture discussion

  • [30m] Meet with CEO

Note: you are welcome to request additional conversations with anyone you would like to meet and we will accommodate as best we can.

Not sure if this is you?

We want a diverse, global team with a broad range of experience and perspectives. If this job sounds great, but you’re not sure if you qualify, we encourage you to reach out to us at [email protected] or subscribe to our career newsletter by emailing "Subscribe" to [email protected]. We carefully consider every application and will either move forward with you, find another team that might be a better fit, keep in touch for future opportunities, or thank you for your time.

Learn More About Us

At Roboflow, we believe great ideas come from everywhere—and everyone. We’re proud to be an Equal Opportunity Employer committed to building a diverse and inclusive team. We consider all qualified applicants regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, age, veteran status, or any other legally protected characteristics.

About Roboflow

Platform for building and deploying custom computer vision models.

Year founded
2019
Employees
108
Organization type
Private
Latest investment
Raised $40.00M Series B (2024) — led by Google Ventures
Headquarters
US

Similar jobs

Machine Learning Engineer roles near New York City, New York
9h
Save
Mark Applied
Hide
Machine Learning Engineer - Satellite Capacity Optimization & Planning
Carlsbad or San Jose or San Francisco or New York City
$141k-$222k/yr OnsiteFull Time
Viasat
ViasatNASDAQ: VSAT: Provider of global satellite-based connectivity and secure communication solutions.
7+ YOERequires 7+ years in ML or optimization, optimization techniques, production cloud ML systems, ML frameworks, geospatial visualization, containerized development, SQL, RESTful APIs, and up to 10% travel.
AWS, GCP, TensorFlow, PyTorch, SQL, RESTful APIs, Airflow, Athena, BigQuery, ECS, Batch
9h
Save
Mark Applied
Hide
Machine Learning Engineer
Toronto or New York City
$150k-$250k/yr HybridFull Time
GPTZero
GPTZero: AI-powered platform for detecting and verifying machine-generated text.
3+ YOERequires 3+ years with PyTorch/Transformers, cutting-edge machine learning experience, software engineering ability, project ownership, leadership, and authorization to work in Canada or the US.
PyTorch, Transformers, AI agents, retrieval-augmented language models, open-source
1d
Save
Mark Applied
Hide
Machine Learning Engineer
Chicago or Washington or Tallahassee or Sarasota or Hartford or Nashville or Costa Mesa or San Jose or Atlanta or Boston or Burlington or Cleveland or Columbus or Dallas or Denver or Fort Lauderdale or Fort Wayne or Grand Rapids or Indianapolis or Knoxville or Lexington or Livingston or Plano or Louisville or Los Angeles or Miami or The Woodlands or New York City or Oakbrook Terrace or Sacramento or San Francisco or South Bend or Springfield or Tampa or Houston or Austin or Charlotte
$62k-$100k/yr OnsiteFull Time
Crowe
Crowe: Global professional services firm providing audit, tax, and consulting.
Expert Python and Git skills; ability to develop, test, debug, and document production AI solutions, navigate ambiguity, communicate technical details, and support ethical and security reviews.
Python, Git, ChatGPT Enterprise, Microsoft Copilot, Azure AI Foundry
1d
Save
Mark Applied
Hide
Staff Machine Learning Engineer, Generative AI Modeling and Inference
Los Angeles or Seattle or Palo Alto or New York City or Bellevue
$195k-$343k/yr OnsiteFull Time
Snap
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
8+ YOEBachelor's degree or equivalent experience and 8+ years of post-bachelor's ML experience, or advanced degree with equivalent experience. Requires computer vision or generative modeling and ML framework experience.
TensorFlow, PyTorch, JAX, MLX, scikit-learn, GPU, CPU, NPU, RSUs
1d
Save
Mark Applied
Hide
Staff Machine Learning Engineer, Generative AI Modeling and Inference
Los Angeles or Seattle or Palo Alto or New York City or Bellevue
$229k-$343k/yr OnsiteFull Time
Snap
SnapNYSE: SNAP: Provides visual messaging software and augmented reality wearable devices.
8+ YOEBachelor's degree or equivalent experience and 8+ years of post-bachelor's machine learning experience, with expertise in generative modeling, computer vision, deep learning, and ML frameworks.
TensorFlow, PyTorch, JAX, MLX, scikit-learn
1d
Save
Mark Applied
Hide
Compliance, Machine Learning Engineer, New York, Vice President
New York City, New York, United States
$130k-$250k/yr OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
10+ YOEBachelor’s or master’s in computer science or related field; 10+ years building scalable machine learning systems; strong coding and systems fundamentals; expertise in Python, PySpark, distributed technologies, and ML toolkits.
Python, PySpark, Scala, Iceberg, HDFS, Avro, Parquet, AWS, GCP, TensorFlow, PyTorch, Scikit-learn, Hugging Face
1d
Save
Mark Applied
Hide
Machine Learning Engineer, Sponsored Products Off-Search Sourcing and Relevance
New York, New York, United States
$158k-$214k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOERequires 3+ years of professional software development, 2+ years designing or architecting scalable systems, and programming experience. Bachelor's degree and machine learning experience preferred.
1d
Save
Mark Applied
Hide
Senior Machine Learning Engineer, TikTok BRIC Community Health
San Jose or Los Angeles or Singapore or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
$162k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOEMaster's degree or above in a relevant technical field and 2+ years of machine learning experience. Requires strong software engineering, machine learning, Python or Java/C++/Go, and Spark, Hadoop, or Hive experience.
Python, Java, C++, Go, Spark, Hadoop, Hive, LLMs