📋 External Recruiting Agencies

Bright Vision Technologies is a staffing and IT consulting firm that recruits talent for clients and provides H-1B sponsorship services, as evidenced by its own self-description and job postings mentioning direct clients.

This company was flagged and excluded from default search results. Proceed with caution.

Bright Vision Technologies
Posted 1w ago

Machine Learning Data Engineer

Bright Vision Technologies
United States
$80k-$100k/yrRemoteFull Time
Responsibilities
  • building data systems
  • operating pipelines
  • delivering datasets
Requirements
  • Bachelor’s or Master’s in Computer Science or related field
  • 6+ years of data engineering experience supporting ML/AI workloads
  • Python and JVM or systems language proficiency
  • Spark, Ray, or Beam experience
Technical tools mentioned
PythonSparkRayBeamCI/CD

Job description

Machine Learning Data Engineer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Machine Learning Data Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $80,000–$100,000 Annually
Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary
We are seeking an Machine Learning Data Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.

Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of data engineering experience, with significant work supporting ML or AI workloads.
  • Strong proficiency in Python and at least one JVM or systems language.
  • Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.
  • Hands-on experience operating petabyte-scale storage and pipeline systems.
  • Strong understanding of distributed systems, data modeling, and storage formats.
  • Experience with dataset versioning, lineage, and reproducibility for ML workflows.
  • Familiarity with high-throughput data loading for accelerator-based training.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
  • Experience with multimodal datasets at large scale.
  • Familiarity with data quality tooling and dataset evaluation methodology.
  • Exposure to privacy-preserving data systems and regulated data handling.
  • Open-source contributions to data infrastructure projects.
  • Experience supporting frontier model training pipelines.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected].
Bright Vision Technologies is an Equal Opportunity Employer.
 

Similar jobs

Machine Learning Data Engineer roles
1w
Save
Mark Applied
Hide
Machine Learning Data Engineer (AIM2)
Cambridge, Massachusetts, United States
$106k-$177k/yr HybridFull Time
Pfizer
PfizerNYSE: PFE: Develops and manufactures vaccines and medicines for global healthcare
2+ YOEPhD in a related technical discipline, or master's degree plus 2 years' experience. Requires Nextflow, NGS, Python, data integration, bioinformatics, and cross-functional collaboration experience.
Nextflow, CI/CD, Python, Claude Code
3w
Save
Mark Applied
Hide
Machine Learning Data Engineer (DataOps), Materra
Mountain View, California, United States
$166k-$244k/yr OnsiteFull Time
X, The Moonshot Factory
X, The Moonshot FactoryNASDAQ: GOOGL: Develops early-stage technologies and experimental projects.
3+ YOE3+ years building scalable data pipelines, strong Python and SQL skills, experience with data validation, dataset versioning, and ML data lifecycle for annotation and training.
Python, Pandas, NumPy, SQL, BigQuery, Cloud Storage, Dataflow, Dataproc, Vertex AI Data Pipelines, Google Cloud Composer, Apache Airflow, Prefect, Dagster, Great Expectations, DVC, TFX, Data Validation
1mo
Save
Mark Applied
Hide
Senior Data & Machine Learning Engineer
Malvern, Pennsylvania, United States
HybridFull Time
Akuvo
Akuvo: Provides AI-powered collections and credit risk software for financial institutions.
6+ YOE6+ years building and deploying production ML models; strong Python and SQL; experience with scikit-learn,XGBoost,LightGBM; Azure ML,Databricks,MLflow experience; model validation, monitoring, and governance.
Python, SQL, scikit-learn, XGBoost, LightGBM, PyTorch, TensorFlow, Azure Machine Learning, Databricks, MLflow, Microsoft Fabric, OneLake, Azure Synapse, Azure DevOps
1mo
Save
Mark Applied
Hide
Senior Machine Learning Data Engineer
Sunnyvale, California, United States
$150k-$240k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
5+ YOE5+ years experience building ML data pipelines and infrastructure, familiarity with ML infrastructure, data-centric AI, GPUs, microservices and databases; U.S. citizenship and ability to obtain security clearance required.
React, TypeScript, Python, Golang, Docker, Kubernetes, OpenSearch, Postgres, GPUs, LLMs, VLMs
3mo
Save
Mark Applied
Hide
Artificial Intelligence/Machine Learning Data Engineer
Charlotte or Morristown or Jersey City or Cleveland or California or Colorado or District of Columbia or Illinois or Maryland or Massachusetts or Minnesota or New York or New Jersey or Washington
$44-$54/hr HybridFull Time
Accenture
AccentureNYSE: ACN: Global provider of management consulting and technology services.
5+ YOEMinimum 5 years in data/ML/AI or software engineering; strong Python and SQL; experience with LLMs, agent frameworks (LangChain, Semantic Kernel, LlamaIndex, AutoGen), ML deployment and data pipelines; Associate's degree required.
MLflow, Azure ML, SageMaker, Kubeflow, TensorFlow, PyTorch, Scikit-Learn, Spark, Databricks, FAISS, Pinecone, Chroma, Milvus, Redis, CI/CD, Docker, Kubernetes, Azure, AWS, GCP, Kafka, EventHub, Kinesis, LangChain, Semantic Kernel, LlamaIndex, AutoGen, Python, SQL, Scala, Java, Go
5mo
Save
Mark Applied
Hide
Machine Learning Data Engineer
Austin, Texas, United States
OnsiteFull Time
Allen Control Systems
Allen Control Systems: Develops autonomous robotic weapon systems to counter drone threats.
3+ YOE3+ years data engineering; AWS; Python; SQL; Linux; data pipelines; computer vision data; strong communication.
AWS, Python, SQL, Linux, PyTorch, Unreal Engine
6mo
Save
Mark Applied
Hide
Machine Learning Data Engineer
New York City, New York, United States
$120k-$190k/yr HybridFull Time
Oden Technologies
Oden Technologies: Provides AI-driven production analytics and recommendations for manufacturers.
1+ YOE1-3 years ML/Data Engineer experience; strong Python; distributed processing tools; ML workflows; SQL; GCP; experience with unstructured data and Generative AI; based in NY metro area; willing to office visit.
Apache Beam, Apache Spark, MLFlow, Vertex AI, SageMaker, Python, SQL, GCP
4w
Save
Mark Applied
Hide
Machine Learning Ops Data Engineer
Southlake or Austin or Phoenix
$102k-$185k/yr OnsiteFull Time
Charles Schwab
Charles SchwabNYSE: SCHW: Financial services, brokerage, and investment management provider.
8+ YOE2+ Mgmt8+ years data/software engineering with 2+ years technical leadership; expert GCP (BigQuery,Vertex AI,GCS,Dataflow,Pub/Sub,Cloud Run,GKE,Composer,Airflow), Python, SQL, CI/CD, containerization, security, observability; experience productionizing ML.
BigQuery, Vertex AI, GCS, Dataflow, Pub/Sub, Cloud Run, GKE, Composer, Airflow, IAM, Cloud Monitoring, Cloud Logging, Python, SQL, Docker, Git