Acclinate
Posted 1w ago

Lead Data Engineer

Acclinate
United States
$79k-$116k/yrRemoteFull Time
Responsibilities
  • building pipelines
  • maintaining integrations
  • optimizing storage
Requirements
  • Requires 6+ years building production data pipelines
  • 2+ years owning a data platform
  • Cloud warehouse experience
  • Regulated health-data experience, and a technical bachelor's degree or equivalent
Technical tools mentioned
SQLBigQueryGoogle Cloud PlatformIAMPythonn8nHubSpotOperations Hub EnterpriseMetabaseVertexAIMetaInstagramYouTubeGoogle AnalyticsTypeformClaudeGit

Job description

Acclinate is seeking a hands-on Data & Analytics Lead to establish a connected, reliable data environment that enables insight-driven decision making, supports business development, and demonstrates measurable impact across our clinical trial engagement efforts. This role will unify currently distributed data sources, build a scalable and practical data infrastructure, and translate complex healthcare, community, and operational data into actionable intelligence for internal teams, partners, and future stakeholders. The position combines technical execution with analytical leadership to ensure Acclinate’s data assets are trusted, usable, and aligned with long-term growth.

Description


As Acclinate's Lead Data Engineer, you will design, build, and operate the data infrastructure that powers the company's analytics, CRM, marketing, and machine learning initiatives. This is a hands-on role: you will be equally responsible for long-term architectural strategy and for the day-to-day engineering, pipeline maintenance, and reporting work that keeps Acclinate's data ecosystem running. You'll work across our event-driven architecture, CRM data platform, BI tooling, and production ML systems, partnering closely with the data lead, engineering, marketing, community, and customer success teams.


Key Responsibilities

  • Design Event-Driven Data Models and Tables: Architect logical and physical data models, creating necessary data tables and schemas to effectively capture, store, and make accessible vast amounts of data generated by event streams. This includes defining how events are transformed into structured data suitable for analysis.
  • Orchestrate Data Integration for Advanced Analytics: Design integration strategies to combine data from event-driven systems with existing data sources, ensuring a unified and consistent view for sophisticated analyses, including data warehousing and data mart design.
  • Co-develop Acclinate's Data Strategy: Collaborate with appropriate unit leads and team members to define Acclinate's long-term vision for data, aligning it with business objectives, and creating a roadmap for implementation across both traditional and event-driven data.
  • Contribute to Data Governance Frameworks: Collaborate with appropriate unit leads and team members to define data ownership, stewardship, quality standards, metadata management, and data lineage tracking for all data, including event data.
  • Ensure Data Security and Compliance: Partner with engineering and legal teams to implement data security controls, access controls, and ensure adherence to relevant data privacy regulations (including PII and sensitive health information) for all data types.
  • Optimize Data Storage and Performance: Recommend and implement appropriate data storage technologies and indexing strategies to ensure efficient data retrieval and analytical query performance, especially for high-volume event data.
  • Evaluate and Recommend New Technologies: Stay abreast of industry trends and propose new data technologies and tools that can benefit Acclinate's data platform.

Hands-On Data Operations & Execution:

  • Oversee Data Integration and ETL/ELT Processes: Build and maintain efficient, reliable data pipelines (e.g., in n8n or equivalent tooling) that transform event streams and other source data into a unified data environment such as a data warehouse or data lake, primarily on GCP/BigQuery.
  • Own CRM Data (HubSpot): Manage HubSpot property and schema design, maintain the HubSpot-to-BigQuery integration (Operations Hub Enterprise), and perform ongoing data cleanup and reconciliation.
  • Build and Maintain BI/Analytics Enablement: Build and maintain Metabase dashboards and a shared metrics library; produce recurring monthly and quarterly reports for Customer Success, Community, Marketing, and project reporting needs.
  • Support Production Machine Learning (PPI Model): Own the ongoing data needs of a live production ML model, including feature engineering, data preparation and imputation, and managing model data in VertexAI; coordinate deployment-related data inputs with the product and engineering teams.
  • Manage Marketing, Social, and Survey Data Pipelines: Build and maintain daily ingestion pipelines from Meta, Instagram, YouTube, Google Analytics, and Typeform into BigQuery.
  • Manage External Data Partnerships: Own tokenized- and identified-data workflows and vendor relationships to support secure, compliant data sharing.
  • Build AI Tooling: Develop and maintain Claude data skills for non-technical teams, including appropriate validation and guardrails.
  • Maintain Data Documentation and Handover Materials: Create and maintain data dictionaries, runbooks, and onboarding materials to support knowledge transfer and business continuity.
  • Apply Engineering Best Practices: Manage all pipeline and transformation code in version control (git), participate in code review, and maintain automated tests and data quality checks on business-critical pipelines.
  • Own Data Quality and Reliability: Define and monitor freshness, completeness, and accuracy checks on critical tables; triage pipeline failures and data incidents, and communicate impact and resolution to affected teams.
  • Manage Platform Cost: Monitor and optimize BigQuery and GCP spend, including query cost, storage tiering, and scheduled job efficiency.
  • Support Compliance Operations: Maintain least-privilege access controls and periodic access reviews on data assets, apply Acclinate’s de-identification standard to externally shared datasets, and produce evidence artifacts (access logs, data lineage, data flow documentation) for audits and client security reviews.


Qualifications

  • 6+ years building and operating production data pipelines, including at least 2 years owning a data platform end to end.
  • Demonstrated ownership of a cloud data warehouse in production. BigQuery preferred; Snowflake, Redshift, or Databricks considered.
  • Experience handling regulated health data (HIPAA/PHI) or comparable regulated data.
  • Bachelor’s degree in a technical field, or equivalent practical experience.


Required Tools and Technical Skills

  • SQL and BigQuery (primary data platform); Google Cloud Platform (IAM, storage, jobs)
  • Python for data preparation, imputation, and analysis notebooks
  • Pipeline and automation tooling (e.g., n8n or equivalent) and scheduled job management
  • HubSpot CRM data model and its BigQuery integration (Operations Hub Enterprise)
  • BI tooling, particularly Metabase
  • ML/MLOps exposure: VertexAI and feature pipelines (supporting production models, not research-grade model development)
  • Health-data privacy and compliance awareness (HIPAA/PHI)
  • Practical AI tooling experience, including building and validating LLM-based "skills"/agents


Nice to Haves

  • HubSpot Ops Hub data model and its BigQuery sync; VertexAI and production ML feature pipelines; building and validating LLM-based agents or “skills”. Strong candidates on the required list will be considered without these. 


Employee Benefits

We offer a comprehensive compensation and benefits package designed to attract and retain top talent. Here's what you can expect:

  • Competitive salary commensurate with your experience and qualifications.
  • Access to health, dental, and vision insurance plans, with the company covering half of the premium costs.
  • Access to 401(k)-retirement plan to help you secure your financial future.
  • Enjoy 22 paid holidays throughout the year.
  • Receive an educational stipend to support your ongoing professional development.
  • A technology allowance to keep you equipped with the tools you need.
  • Generous paid time off to recharge and maintain work-life balance.



About the Company


Acclinate is the catalyst for health equity. We use trust, technology, and community engagement to help the healthcare ecosystem connect with those they aim to serve. Our approach ensures innovation is inclusive, empowering communities to take actions for better health.

About Acclinate

Platform for increasing diversity in clinical trial participation

Year founded
2020
Employees
50
Organization type
Private
Latest investment
Raised $7.00M Series A (2024) — led by Cencora Ventures
Subsidiaries
Headquarters
US

Similar jobs

Data Engineer roles
4h
Save
Mark Applied
Hide
Senior Data Engineer
Chicago or United States
RemoteFull Time
Tiger Analytics
Tiger Analytics: Provides AI and advanced analytics consulting for global enterprises.
8+ YOERequires 8+ years in data engineering, strong AWS, Databricks, Spark, SQL, ETL/ELT, and Airflow experience, plus pharmaceutical data, modeling, lakehouse, analytics, and troubleshooting expertise.
Amazon S3, AWS Glue, AWS Lambda, Amazon Redshift, Databricks, Apache Spark, SQL, Apache Airflow, ETL, ELT, LLM, RAG
4h
Save
Mark Applied
Hide
Staff Data Engineer - Core Data Pipelines
Detroit, Michigan, United States
OnsiteFull Time
KODE Labs
KODE Labs: Provides a cloud-based operating system for smart buildings.
Staff-level data engineering experience, scalable pipeline design, cross-team technical leadership, SQL and Python proficiency, and experience with high-volume data, data modeling, quality, lineage, and observability.
SQL, Python, Kafka, Flink, Spark, Airflow, dbt, BigQuery, ClickHouse, PostgreSQL
4h
Save
Mark Applied
Hide
Technical Lead-Data Engineer
Irvine, California, United States
$130k-$200k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
10+ YOE2+ MgmtRequires 10+ years in data engineering/backend development, 2+ years leading or mentoring, and expertise in Databricks, dbt, Airflow, Python, SQL, and a cloud platform.
Databricks, dbt, Apache Airflow, PySpark, Delta Live Tables (DLT), Spark SQL, GitHub Actions, Azure DevOps, JIRA, Python, SQL, AWS, Azure, GCP, Delta Lake, Databricks Unity Catalog
4h
Save
Mark Applied
Hide
Senior Data Engineer - 26206
United States
OnsiteFull Time
Enverus
Enverus: Software and analytics for the global energy industry.
Bachelor's degree in computer science or related quantitative field, deep production software or data-systems experience, large-scale distributed systems expertise, strong software engineering and SQL skills, and technical initiative leadership.
SQL
4h
Save
Mark Applied
Hide
Sr. Data Engineer, Team Lead, Store Launch
St. Louis, Missouri, United States
OnsiteFull Time
DriveCentric
DriveCentric: AI-powered CRM platform built for automotive dealerships.
5+ YOE2+ MgmtRequires 5+ years SQL/database development, 3+ years T-SQL and SQL Server, 3+ years database solutions, ETL and data import expertise, programming or scripting, and 2+ years people management.
SQL, T-SQL, SQL Server, C#, PowerShell, Python, AWS, S3, EC2, RDS, Snowflake, Databricks, dbt, Pluralsight
4h
Save
Mark Applied
Hide
Data Engineer
United States or Reston
RemoteFull Time
Resonate
Resonate: AI-powered platform providing consumer intelligence and marketing data analytics.
5+ YOERequires 5+ years in software or data engineering, 3+ years with Spark and Scala, multi-terabyte or petabyte Spark tuning, relational databases, cloud big data stacks, testing, architecture, and production operations.
Apache Spark, Scala, Amazon Web Services (AWS), Amazon EMR, Amazon S3, Snowflake, Grafana, Apache Kafka, Hadoop, Elastic Stack, Docker, Amazon Lambda
5h
Save
Mark Applied
Hide
Staff Data Engineer, Individual Contributor | Data & Machine Learning
Arizona or Colorado or Florida or Georgia or Idaho or Illinois or Kansas or Massachusetts or Michigan or Minnesota or Missouri or New Hampshire or New York or North Carolina or Ohio or Oregon or Pennsylvania or South Carolina or Tennessee or Texas or Utah or Virginia or Washington or Wisconsin
$180k-$200k/yr RemoteFull Time
MedBridge
MedBridge: Provides healthcare education and patient engagement software solutions.
Deep cloud data-platform experience; advanced Snowflake, Python, and SQL skills; production pipeline, security, governance, CI/CD, and cross-functional technical leadership experience; technical or quantitative degree or equivalent capability.
Snowflake, Python, SQL, Snowpark, Model Registry, ML Jobs, Openflow, dbt, Terraform, EHR, EMR
6h
Save
Mark Applied
Hide
Senior Data Engineer (GCP • Python • Iceberg • Delta Lake • Kafka • Snowflake • Databricks)
United States
$120k-$180k/yr RemoteFull Time
Railroad19
Railroad19: Builds custom enterprise software and cloud solutions for clients.
6+ YOERequires 6+ years of enterprise Python and Spark experience, advanced GCP BigQuery, Iceberg, Delta Lake, Delta Sharing, Kafka/CDC, Snowflake Horizon, Databricks Unity Catalog, and data lake architecture expertise.
Google Cloud Platform (GCP), Python, Apache Spark, Google Cloud Storage (GCS), BigQuery, Apache Iceberg, Delta Lake, Iceberg UniForm, Delta Sharing, Kafka, Snowflake, Snowflake Horizon Catalog, Databricks, Databricks Unity Catalog, LookML, Claude Code