Bright Vision Technologies
Posted 2d ago

Infrastructure Reliability Engineer

Bright Vision Technologies
United States
$125k-$170k/yrRemoteFull Time, Contract
Responsibilities
  • designing pipelines
  • building ingestion
  • optimizing performance
Requirements
  • Bachelor’s or Master’s in Computer Science or related field
  • 6+ years of data engineering experience supporting ML/AI
  • Python and JVM or systems language proficiency, and Spark, Ray, or Beam experience
Technical tools mentioned
PythonSparkRayBeamCI/CD

Job description

Infrastructure Reliability Engineer  – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Infrastructure Reliability Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $125,000–$170,000 Annually
Experience Required: 6+ years

Job Summary
We are seeking an Infrastructure Reliability Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.

Key Responsibilities
  • Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement workflows.
  • Build ingestion systems for diverse modalities including text, image, audio, video, and structured signals.
  • Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale.
  • Develop dataset versioning, lineage, and provenance tracking systems suitable for reproducible training.
  • Build high-throughput data loading systems that maximize GPU utilization during training.
  • Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement systems.
  • Design storage architectures balancing cost, throughput, and latency across data tiers.
  • Build evaluation dataset construction pipelines with strict integrity and contamination controls.
  • Implement data privacy, redaction, and consent enforcement throughout the pipeline.
  • Collaborate with ML researchers and engineers to align data systems with model development needs.
  • Drive observability of data quality, drift, and pipeline health across the AI data estate.
  • Optimize cost and performance through compression, format selection, and caching strategies.
  • Document data systems, schemas, and operational procedures for broad internal use.
  • Stay current with AI data infrastructure research and emerging open-source tools.

Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of data engineering experience, with significant work supporting ML or AI workloads.
  • Strong proficiency in Python and at least one JVM or systems language.
  • Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.
  • Hands-on experience operating petabyte-scale storage and pipeline systems.
  • Strong understanding of distributed systems, data modeling, and storage formats.
  • Experience with dataset versioning, lineage, and reproducibility for ML workflows.
  • Familiarity with high-throughput data loading for accelerator-based training.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.

Preferred Qualifications
  • Experience with multimodal datasets at large scale.
  • Familiarity with data quality tooling and dataset evaluation methodology.
  • Exposure to privacy-preserving data systems and regulated data handling.
  • Open-source contributions to data infrastructure projects.
  • Experience supporting frontier model training pipelines.

How to Apply

Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908)676-4399. Learn more about Bright Vision Technologies at www.bvteck.com.

Bright Vision Technologies is an Equal Opportunity Employer.

About Bright Vision Technologies

AI-powered enterprise automation and software development firm.

Similar jobs

Infrastructure Reliability Engineer roles
2w
Save
Mark Applied
Hide
Sr. Infrastructure Reliability Engineer, Infrastructure Reliability & Quality
Herndon, Virginia, United States
$137k-$185k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
6+ YOERequires a bachelor's degree in engineering or related field and 6+ years in data center or mission-critical facilities engineering. Expertise in reliability risk analysis, statistical methods, equipment, quality, root cause analysis, and vendor management.
Physics-of-Failure, reliability block diagram, statistical modeling, data analytics
3mo
Save
Mark Applied
Hide
Infrastructure Reliability Engineer
Manassas or Sterling or Portland or Chicago or Dallas Fort Worth
OnsiteFull Time
STACK Infrastructure
STACK Infrastructure: Private global data-center developer and operator providing colocation, build-to-suit, and powered-shell infrastructure to hyperscale and enterprise customers.
5+ YOE5–8 years in critical infrastructure; strong fluency in electrical systems; RCA/forensic troubleshooting; bachelor’s in engineering or equivalent.
Power distribution equipment, Waveform analysis, Fault analysis tools
4mo
Save
Mark Applied
Hide
Staff Infrastructure Reliability Engineer - Database & Storage
Seattle or San Francisco or Detroit or United States
$180k-$279k/yr HybridFull Time
Rocket Homes
Rocket Homes: Technology-driven real estate service provider connecting home buyers and sellers with listings and real estate agents.
7+ YOE7+ years AWS/cloud infra; 5+ years PostgreSQL/AWS services; Linux admin/scripting; mentoring; infrastructure as code and security; AI code generation tools; on-call readiness.
AWS, PostgreSQL, Aurora/RDS, S3, ElastiCache, OpenSearch, DynamoDB, Linux, Python, Infrastructure as Code, Security practices, AI code generation tools
5mo
Save
Mark Applied
Hide
Software Engineer, Infrastructure Reliability
San Francisco, California, United States
$255k-$405k/yr OnsiteFull Time
OpenAI
OpenAI: AI research and deployment focused on beneficial AGI.
4+ YOE4+ years relevant experience; 2+ years leading large-scale projects; expert in distributed systems, cloud infrastructure, and observability tools.
Kubernetes, Terraform, AWS, GCP, Azure, Datadog, Prometheus, Grafana, Splunk, ELK, CI/CD, Linux, Observability, Service mesh
8mo
Save
Mark Applied
Hide
Infrastructure and Reliability Engineer - Developer Platform
San Francisco, California, United States
$180k-$280k/yr OnsiteFull Time
TypeSafe AI
TypeSafe AI: Private frontier AI lab building reliable, general AI systems for real-world automation and decision-making.
5+ YOE5+ years software engineering; 3+ years infra/backend; Kubernetes, cloud providers, AWS; ML orchestration experience; strong debugging under pressure; team collaboration and ownership.
Python, TypeScript, Next.js, Tailwind CSS, Kubernetes, Claude Code, Cursor
1w
Save
Mark Applied
Hide
Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere)
San Francisco or Denver or Austin or Jacksonville or Bridgeport or Seattle or Boston or New York City
$126k-$205k/yr RemoteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Global cybersecurity platform providing network, cloud, and AI-driven security solutions.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
Go, Java, Python, Rust, AWS, GCP, Kubernetes, Linux, Terraform, Cursor, Claude
3w
Save
Mark Applied
Hide
Senior Site Reliability and Infrastructure Engineer
New York City, New York, United States
$160k-$220k/yr HybridFull Time
Treeswift
Treeswift: Robotics and physical-AI helping electric utilities’ field crews get more work done.
7+ YOE7+ years experience in observability, SRE, infrastructure or DevOps; hands-on Terraform, Kubernetes, Linux, CI/CD; experience with Airflow-style pipelines and cloud services; strong debugging and communication skills.
Apache Airflow, Astronomer, AWS, Kubernetes, S3, SQS, Lambda, Step Functions, ECS, ECR, Astronomer CLI, Terraform, DuploCloud, Linux