🏢 Offshore Consulting Shops

Coderoad is an IT services and staff augmentation firm that provides outsourced software development teams and engineering talent to other companies.

This company was flagged and excluded from default search results. Proceed with caution.

C
Posted 6mo ago

Machine Learning Operations Engineer (MLOps)

CodeRoad
South America
RemoteContract
Responsibilities
  • Design deployment
  • Manage cloud infra
  • Automate CI/CD
Requirements
  • 4+ years in MLOps/DevOps with AI/ML deployment experience
  • Cloud (GCP or AWS)
  • Python
  • Docker/Kubernetes
  • Security and ethics in AI
Technical tools mentioned
PyTorchLangraphCrewAIN8NDockerKubernetesTerraformPythonGitPrometheusGrafanaLantraceAgentOpsVertex AICloud FunctionsSageMakerS3BigQueryRedshiftLambda

Job description

Machine Learning Operations (MLOps) Engineer

The Team

At Coderoad, we're more than just a software development company—we're your gateway to the global tech world. Whether you're looking to skill up or level up your career, we offer the challenges you’ve been searching for.

We provide end-to-end software development services and give you the opportunity to work on exciting, real-world projects in a supportive environment. Whether it's staff augmentation, dedicated IT teams, or general software engineering, we have opportunities for everyone to challenge themselves and take their career to the next level!

Position Location - Latam (Remote).
Time Zone Requirements - This team operates on the East/West Coast time zones.

About the Role

We are seeking a skilled and innovative Machine Learning Operations (MLOps) Engineer with a focus on Agentic AI to design, deploy, and maintain scalable, robust, and ethical autonomous AI systems. The ideal candidate will combine deep expertise in modern MLOps practices with a solid understanding of agentic AI principles, enabling the seamless integration, monitoring, and optimization of AI models that exhibit autonomous decision-making and adaptability. You will be a key contributor in our cross-functional teams, ensuring our agentic AI solutions are reliable, efficient, and aligned with our business goals and ethical standards.

Key Responsibilities

  • Model Deployment & Integration: Design and implement scalable, secure, and production-grade pipelines for deploying agentic AI models. Focus on seamless integration with existing systems and enable real-time adaptability for autonomous decision-making.

  • Cloud Infrastructure Management: Build and maintain robust cloud infrastructure on Google Cloud Platform (GCP) or Amazon Web Services (AWS) for the entire AI lifecycle. Leverage services like GCP's Vertex AI and Cloud Functions, or their AWS equivalents such as Amazon SageMaker, and AWS Lambda, to create efficient and resilient environments.

  • Automation & CI/CD: Develop and maintain automated workflows for continuous integration, continuous deployment (CI/CD), and continuous training (CT) of agentic AI models. Optimize for performance, scalability, and reliability using CI/CD platforms.

  • Monitoring & Performance Optimization: Implement and manage advanced monitoring systems to track the performance, health, and decision-making accuracy of agentic AI models in production. Utilize specialized tools like Lantrace, AgentOps, or AWS's CloudWatch to detect and resolve issues related to model drift, latency, and bias in real-time.

  • Security & Compliance: Integrate security best practices throughout the MLOps lifecycle. Ensure agentic AI systems adhere to ethical guidelines and regulatory requirements, implementing safeguards for data privacy, bias mitigation, and transparency in autonomous operations.

  • Collaboration: Work closely with AI researchers, data scientists, software engineers, and product teams to align MLOps processes with project goals. Facilitate iterative development and deployment of agentic AI solutions.

  • Data & Model Governance: Establish and enforce robust data and model governance frameworks, ensuring data quality, security, and compliance with industry standards for all agentic AI systems.

Qualifications

  • Experience: 4+ years of experience in MLOps, DevOps, or a related field, with at least 1 year focused on deploying and managing AI/ML models in production. Experience with agentic or autonomous AI systems is highly preferred.

  • Cloud Expertise: (4years)Deep hands-on experience with either Google Cloud Platform (GCP) or Amazon Web Services (AWS). Knowledge of relevant services such as GCP's Vertex AI, Cloud Storage, BigQuery, and Cloud Functions or AWS equivalents like Amazon SageMaker, S3, Redshift, and Lambda.

  • Technical Stack: (1 year or less)Strong knowledge of MLOps tools and frameworks(Pytorch, Langraph, CrewAI, N8N). Proficiency in containerization with Docker and orchestration with Kubernetes.

  • Programming & Scripting: Expertise in Python and familiarity with scripting for automation (e.g., Bash, Terraform). Strong experience with version control systems, particularly Git.

  • Monitoring & Analytics: Hands-on experience with modern monitoring tools like Lantrace, AgentOps, Prometheus, or AWS's CloudWatch and Grafana. Proven ability to track model performance, data drift, and system health in a production environment.

  • Security Mindset: A strong understanding of security principles related to cloud and MLOps, including Identity and Access Management (IAM), data encryption, and secure pipeline design.

  • Ethical AI Knowledge: Understanding of ethical AI principles, including bias detection, explainability, and compliance with regulations like GDPR or other relevant standards.

  • Collaboration & Communication: Strong interpersonal and communication skills, with the ability to work effectively in cross-functional teams and explain technical concepts clearly to diverse stakeholders.

Education: Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field. Advanced degrees or certifications in MLOps, AI/ML, or cloud technologies are highly valued.

What you’ll love:
  • 100% Remote

  • Contractor position available for Latin American candidates

  • Holidays Off

  • Paid Time Off

  • Health insurance assistance program.

  • Competitive Pay (USD)

  • Excellent teamwork and work environment

  • Training

Similar jobs

Machine Learning Operations Engineer roles
13h
Save
Mark Applied
Hide
(Senior) Machine Learning Operations Engineer (m/f/d) - REF97172N
Hannover, Lower Saxony, Germany
HybridFull Time, Part Time
Continental
ContinentalXetra: CON: Manufacturer of tires, automotive parts, and industrial rubber products.
Degree in computer science, IT, engineering, mathematics, or related field; several years of MLOps experience; Python, SQL, CI/CD, containers, cloud, ML deployment, monitoring, and Agile expertise.
Python, SQL, GitHub Actions, Docker, Kubernetes, MLflow, SageMaker, AWS, Azure, Grafana, Airflow, GitOps, Scrum, Kanban, LinkedIn Learning, JobRad
1d
Save
Mark Applied
Hide
LEAD MACHINE LEARNING OPS ENGINEER
Solna, Stockholm County, Sweden
OnsiteFull Time
SAS
SAS: Major Scandinavian airline providing passenger and cargo flight services.
5+ YOERequires a master's degree, 5+ years of ML engineering or MLOps experience, Azure production experience, Python, MLflow, Docker, Kubernetes, Terraform, and strong communication skills.
Azure, Azure ML, Azure Databricks, Azure Data Factory, Python, MLflow, Docker, Kubernetes, Terraform, Coding Assistants
1d
Save
Mark Applied
Hide
Software Engineering, Machine Learning Operations
Mountain View, California, United States
$166k-$244k/yr HybridFull Time
Alphabet
AlphabetNASDAQ: GOOGL: Holding providing internet, software, and AI services.
3+ YOEBachelor's or master's degree in computer science, engineering, or related field; 3+ years in software, DevOps, or data engineering; Python, Docker, cloud, Git, and CI/CD experience, including MLOps or ML infrastructure.
Cloud Build, GitHub Actions, Vertex AI, Google Kubernetes Engine (GKE), Vertex AI Pipelines, Docker, Python, Google Cloud Platform (GCP), Amazon Web Services (AWS), Microsoft Azure, Git, Vertex AI Feature Store, Vertex AI Model Registry, BigQuery, Kubeflow, Terraform, TensorFlow, PyTorch, Scikit-learn
1d
Save
Mark Applied
Hide
Software Engineering, Machine Learning Operations, Tapestry
Mountain View, California, United States
$166k-$244k/yr HybridFull Time
X, The Moonshot Factory
X, The Moonshot FactoryNASDAQ: GOOGL: Develops early-stage technologies and experimental projects.
3+ YOEBachelor's or master's degree in computer science, engineering, or related field; 3+ years in software, DevOps, or data engineering; Python, Docker, cloud, Git, CI/CD, and MLOps experience.
Python, Cloud Build, GitHub Actions, Vertex AI, GKE, Vertex AI Pipelines, Docker, GCP, AWS, Azure, Git, Vertex AI Feature Store, Vertex AI Model Registry, BigQuery, Kubeflow, Terraform, TensorFlow, PyTorch, Scikit-learn
4d
Save
Mark Applied
Hide
Senior Machine Learning Operations Engineer
Somerville or Boston
$150k-$200k/yr OnsiteFull Time
AgZen
AgZen: AI-powered precision agriculture technology for optimized crop spraying.
5+ YOEBachelor's or graduate degree in a related field, 5+ years building distributed or ML systems, MLOps and cloud experience, Python and SQL proficiency, ML framework expertise, and strong communication skills.
Python, PyTorch, TensorFlow, SQL, NumPy, Pandas, scikit-learn
4d
Save
Mark Applied
Hide
Machine Learning Operations Engineer (f/m/div.)
Braga, Braga District, Portugal
HybridFull Time
Bosch
Bosch: Global manufacturer of automotive and industrial engineering technology.
3+ YOEMSc or PhD in a relevant field and 3+ years in MLOps, data engineering, DevOps, or similar roles; Python, Bash, Azure, Terraform, GitHub Actions, Kubernetes, and Linux experience required.
Terraform, Python, Bash, Azure, Microsoft Entra ID, GitHub Actions, Kubernetes, Ansible, GNU/Linux
6d
Save
Mark Applied
Hide
ML OPS til udvikling og støtte af Computer Network Exploitation
Søborg, Capital Region of Denmark, Denmark
OnsiteFull Time
Politi
Politi: The national police force of the Kingdom of Denmark.
Technical curiosity, methodological and solution-oriented approach, ability to work with complex technologies, and experience with LLM inference, AI agents, data processing, RAG, DAG generation, and MCP servers.
Large Language Models (LLM), llama.cpp, vLLM, SGLang, TensorRT-LLM, DAG, RAG, MCP
1w
Save
Mark Applied
Hide
Senior Machine Learning Operations Engineer
United States
$170k-$210k/yr RemoteFull Time
Hungryroot
Hungryroot: Provides AI-powered personalized grocery delivery and meal planning.
5+ YOE5+ years in MLOps, ML engineering, or DevOps; strong Python, SQL, and Bash; experience with AWS, Databricks, Spark, MLflow, CI/CD, Docker, APIs, and infrastructure as code.
Python, FastAPI, AWS, Spark, Databricks, MLflow, Git, GitHub Actions, Jenkins, Terraform, Docker, ECS, EKS, Gurobi, OR-Tools, Statsig, Feast, Tecton, Scala, C++, SQL, Bash, Unity Catalog, Databricks Asset Bundles, Databricks Feature Store