Arcadia
Posted 4w ago

Analytics Engineer, Life Sciences Delivery Operations

Arcadia
United States
$175k-$200k/yrRemoteFull Time
Responsibilities
  • authoring models
  • operating pipeline
  • supporting partners
Requirements
  • 5+ years data engineering with production pipelines
  • Dbt and Spark/PySpark
  • Strong SQL and Python
  • Snowflake and AWS S3 experience
  • HIPAA de-identification knowledge
  • CI/CD and customer-facing communication skills
Technical tools mentioned
dbt-SparkPySparkPythonSQLArgo WorkflowsKubernetesAWS S3Apache IcebergAWS GlacierSnowflakeGitGitHubJiraConfluenceClaudeChatGPT

Job description

Why This Role Is Important to Arcadia

Life sciences customers depend on Arcadia's real-world data to power drug development, safety surveillance, and outcomes research. As LS deal volume accelerates, the engineering foundation underneath delivery, i.e. quality, automation, data transformation evolution, and scale must keep pace.

This is a hybrid role at the intersection of data engineering, data analysis, and delivery operations. You'll refactor, scale, own, and operate an automated RWD data delivery pipeline via dbt/AWS architecture, serving as the primary technical point of contact for channel partners.

You write production-grade PySpark and dbt one day and may facilitate a data inquiry the next. You care deeply about both the correctness of the code and the clarity of the answer it produces. You're as comfortable in a GitHub PR as you are in a partner meeting.

This is a foundational engineering role in a growing LS organization. The right person will help build the team as the business scales.

 

What Success Looks Like

In 3 months

  • Deep familiarity with the end-to-end LS pipeline-from ingestion through dbt transformation, de-identification, and delivery-including the current Snowflake-based scripts and what will replace them
  • Ownership of the channel partner data inquiry queue; resolving standard requests independently by leveraging AI agents, closing out in writing and in accordance with SLAs
  • First contribution to the delivery pipeline codebase: a new or refactored dbt model, a PySpark debugging fix, or a validated QC delivery configuration
  • Thorough understanding of the monthly delivery cycle: Argo orchestration, Snowflake execution, manifest generation, Datavant/HealthVerity/IQVIA tokenization, and delivery QC

In 6 months

  • Core delivery endpoint configurations migrated from manual Snowflake runbook to config-as-code; existing channel partners delivered with minimal manual script execution
  • Contributing increasingly receptive metrics toward a data quality scorecard, tracking pipeline health, completeness, and refresh SLAs across all channel partners
  • PHI de-identification compliance implementation process owned end-to-end, with clear documentation of rules applied
  • Strong working partnerships established with platform engineering (Data Engineering, TechOps) with clear interfaces and shared standards

In 12 months

  • Monthly delivery cycle runs automatically; manual Snowflake execution eliminated; delivery cycle time reduced
  • Recognized internally as the technical authority on the LS data engineering architecture and delivery pipeline
  • Test suites, acceptance criteria, and release documentation authored for all major pipeline changes
  • Potentially beginning to mentor a junior team member as the LS delivery organization grows


Why This Role Is Important to Arcadia
Life sciences customers depend on Arcadia's real-world data to power drug development, safety surveillance, and outcomes research. As LS deal volume accelerates, the engineering foundation underneath delivery, i.e. quality, automation, data transformation evolution, and scale must keep pace.
This is a hybrid role at the intersection of data engineering, data analysis, and delivery operations. You'll refactor, scale, own, and operate an automated RWD data delivery pipeline via dbt/AWS architecture, serving as the primary technical point of contact for channel partners.
You write production-grade PySpark and dbt one day and may facilitate a data inquiry the next. You care deeply about both the correctness of the code and the clarity of the answer it produces. You're as comfortable in a GitHub PR as you are in a partner meeting.
This is a foundational engineering role in a growing LS organization. The right person will help build the team as the business scales.
 
What Success Looks Like
In 3 months
Deep familiarity with the end-to-end LS pipeline-from ingestion through dbt transformation, de-identification, and delivery-including the current Snowflake-based scripts and what will replace them
Ownership of the channel partner data inquiry queue; resolving standard requests independently by leveraging AI agents, closing out in writing and in accordance with SLAs
First contribution to the delivery pipeline codebase: a new or refactored dbt model, a PySpark debugging fix, or a validated QC delivery configuration
Thorough understanding of the monthly delivery cycle: Argo orchestration, Snowflake execution, manifest generation, Datavant/HealthVerity/IQVIA tokenization, and delivery QC
In 6 months
Core delivery endpoint configurations migrated from manual Snowflake runbook to config-as-code; existing channel partners delivered with minimal manual script execution
Contributing increasingly receptive metrics toward a data quality scorecard, tracking pipeline health, completeness, and refresh SLAs across all channel partners
PHI de-identification compliance implementation process owned end-to-end, with clear documentation of rules applied
Strong working partnerships established with platform engineering (Data Engineering, TechOps) with clear interfaces and shared standards
In 12 months
Monthly delivery cycle runs automatically; manual Snowflake execution eliminated; delivery cycle time reduced
Recognized internally as the technical authority on the LS data engineering architecture and delivery pipeline
Test suites, acceptance criteria, and release documentation authored for all major pipeline changes
Potentially beginning to mentor a junior team member as the LS delivery organization grows


What You'll Be Doing

RWD DATA PIPELINE ENGINEERING

  • Author and maintain dbt models and PySpark transformation jobs, replacing ad-hoc Snowflake scripts with governed, version-controlled, tested code
  • Design and implement delivery endpoint configurations as code-customer, delivery target (Snowflake, S3), cadence, cohort filters, incremental and full historical refresh methods
  • Write production-grade Python and PySpark for data transformation, validation automation, and delivery pipeline components, including customer-specific data models and schema validation logic
  • Configure and maintain AWS S3 delivery paths, Apache Iceberg table structures, and file staging patterns for partner data delivery
  • Partner with platform engineering to build and extend Argo Workflows orchestration for automated delivery execution, eliminating the manual monthly Snowflake runbook
  • Implement and maintain HIPAA de-identification compliance rules in pipeline code in accordance with ED certificates; coordinate certification updates when new data elements or rule changes affect certified products
  • DELIVERY OPERATIONS & DATA QUALITY

    • Coordinate and execute monthly RWD deliveries across all active channel partners: delivery job execution, manifest generation and validation, tokenization workflows, and QC
    • Define and monitor delivery quality metrics: pipeline health, data completeness, referential integrity, refresh SLA tracking, and minimum volume thresholds; manage on-time delivery against a >=95% target
    • Validate dbt model outputs and pipeline changes against expected schema and counts; define acceptance criteria and execute UAT before changes reach production or channel partners
    • DATA INVESTIGATIONS & PARTNER SUPPORT

      • Own the channel partner data inquiry queue-triage, investigate, resolve, and communicate on data questions and discrepancies; you are the primary research contact for channel partners
      • Conduct root cause analysis on anomalies in RWD, claims, and clinical feeds, distinguishing source-level issues from transform-layer failures, and communicate findings clearly in writing
      • Build and maintain repeatable query libraries, data dictionaries, and end-to-end pipeline documentation in Confluence, reducing one-off analytical effort and improving institutional knowledge
      • Build out a knowledge repository of learnings from all data inquiries, and partner with VP to provide a data FAQ for LS customers
      • Contribute to data quality governance: build scorecards and trend analyses that equip leadership with evidence-based positions for engineering and product discussions
      • ENGINEERING PRACTICES & COLLABORATION

        • Follow SDLC best practices: author requirements, write test plans, manage releases, and maintain operating documentation in Confluence
        • Manage engineering work through Jira: clear ticket authoring, acceptance criteria, dependency tracking, and proactive status communication
        • Apply CI/CD practices via GitHub: branch management, PR-based review workflows, dbt model testing, and version control discipline-the version in production always matches what's in Git
        • Partner closely with Arcadia's platform engineering team on pipeline architecture, table design, and data contracts
        • Leverage AI tools (including Claude Code) to accelerate development, automate documentation, generate and verify code, and improve operational throughput
        •  

          Technologies

          • Pipeline & transformation: dbt-Spark, PySpark, Python, SQL
          • Orchestration: Argo Workflows, Kubernetes
          • Storage: AWS S3, Apache Iceberg, AWS Glacier
          • Warehousing: Snowflake (delivery target, Iceberg, analytics)
          • Person De-identification and Tokenization Processes
          • Source control & CI/CD: Git / GitHub, PR-based review workflows
          • Workflow & observability: Jira, Confluence, etc.
          • AI tooling: Claude, ChatGPT


What You'll Bring

Education

  • Bachelor's or Master's degree in Computer Science, Data Science, Statistics, or a related field (or equivalent professional experience)
  • Experience

    • 5+ years of hands-on data engineering experience (production pipelines, dbt, Spark/PySpark, cloud data infrastructure) AND 5+ years of direct experience with life sciences RWD data (claims, EHR, clinical); these disciplines can overlap-5 years total is sufficient if you bring meaningful depth in both
    • Production-grade SQL proficiency in Snowflake or a comparable columnar warehouse: complex joins, CTEs, window functions, incremental patterns – you write this fluently
    • Python and/or PySpark for data transformation: you have written and debugged production Spark jobs, not just automation scripts
    • dbt: hands-on experience authoring models, tests, macros, and yml documentation; familiarity with incremental strategies and model validation
    • AWS S3: practical experience with file staging, delivery paths, bucket structure, and lifecycle management in a data engineering context
    • HIPAA de-identification: working knowledge of Safe Harbor requirements and how they are applied in data pipelines before data leaves your custody
    • SDLC fundamentals: you write requirements, author test plans, manage releases, and document your work – this is not new to you
    • CI/CD and source control: Git/GitHub, PR-based review workflows, branching strategies
    • External customer experience: you have led (or actively presented within) technical data discussions with partner analytics or science teams and can communicate complex data concepts clearly in writing and verbally
    • Self-starter who operates independently in ambiguous, high-growth environments and a natural collaborator when the work calls for it

    •  

      Skills

      • Strong analytical judgment – you can look at a distribution and know when something is wrong
      • Clear communicator – able to translate technical pipeline findings for partners and non-technical stakeholders
      • Builder's mindset – you don't just answer the question in front of you, you eliminate the conditions that caused it
      • Genuine curiosity – AI tooling and motivated to apply it to your own workflows, not just in principle
      • Multi-tasking ability – manage multiple parallel workstreams, triage competing priorities, and communicate proactively on risk


Would Love For You To Have
  • Experience with Argo Workflows or a comparable orchestration platform
  • Familiarity with patient tokenization or record linkage technologies and how they integrate into RWD delivery workflows
  • Experience with HIPAA de-identification certification process and implementation workflow
  • Healthcare data standards: ICD-10, CPT, NDC, LOINC, NPI
  • Experience working at a healthcare data vendor, RWD aggregator, or analytics platform company
  • Comfortable operating in an AI-first environment, using Claude or similar tools to build, verify, and accelerate day-to-day workflows
  • Demonstrated interest in people leadership-you'd be excited to mentor and eventually grow a small team as the LS organization scales


What You'll Get
  • Be at the center of a high-stakes, high-impact engineering RWD delivery pipeline you help create will define how Arcadia delivers RWD to life science partners at scale
  • Become the definitive internal expert on one of the most complex and valuable real-world healthcare datasets in the market, with the autonomy to shape how it is engineered, measured, and delivered
  • Be on the front lines of AI adoption-use cutting-edge tools to accelerate your work and shape how the team operates in an AI-first environment
  • Flexible, fully remote work environment, with resources and support to do your best work
  • Exposure to senior leaders across the entire life science and corporate engineering teams
  • A clear path to grow into a player/manager role as Arcadia's life sciences delivery team scales
  • Become a member of the talented, energized, diverse, and purpose-driven Arcadian community


About Arcadia
Arcadia.io helps innovative providers and payers across the country transform healthcare to reduce cost while improving patient health. We do this by aggregating large amounts of disparate data, applying algorithms to identify opportunities to provide better patient care, and making those opportunities actionable by physicians at the point of care in near-real time. We are passionate about helping our customers drive meaningful outcomes. We are growing fast and have emerged as a market leader in the highly competitive population health management software market and have been recognized by industry analysts KLAS, IDC, Forrester, and Chilmark for our leadership. For a better sense of our brand and products, please explore our website.

Protect Yourself
If you have concerns about the authenticity of a job offer or recruitment-related communication claiming to be from Arcadia, we encourage you to verify by contacting us directly at (781) 202-3600 and select option 3. For more information, visit our website.

This position is responsible for following all Security policies and procedures in order to protect all PHI under Arcadia's custodianship as well as Arcadia Intellectual Properties.  For any security-specific roles, the responsibilities would be further defined by the hiring manager.

About Arcadia

Provides a data analytics platform for healthcare organizations.

Year founded
2002
Employees
450
Organization type
Private
Latest investment
Raised $125.00M Debt Financing (2023) — led by Vista Credit Partners
Subsidiaries
Headquarters
US

Similar jobs

Analytics Engineer roles
15h
Save
Mark Applied
Hide
Lead Analytics Engineer
United States
$130k-$150k/yr RemoteFull Time
Innodata
InnodataNASDAQ: INOD: A global providing AI data engineering and services.
9+ YOERequires 9+ years in analytics, business intelligence, or analytics engineering; 3+ years at senior level; advanced SQL, Airflow, data modeling, Python, visualization, metric governance, and English communication.
SQL, Presto, Trino, Hive, Spark SQL, Snowflake, BigQuery, Redshift, Airflow, Dagster, Prefect, Python, pandas, PySpark, dbt, Tableau, Apache Superset, Looker, Power BI, Mode
18h
Save
Mark Applied
Hide
Analytics Engineer
United States
RemoteFull Time
Hipp Health
Hipp Health: AI-native operating system for ambulatory healthcare practices
3+ YOERequires 3+ years in analytics, data engineering, or data analytics; production data pipelines and models; advanced SQL; BI tools; Git; data warehousing; and software development best practices.
Git, Omni, Tableau, Microsoft Power BI, Looker, SQL, DBT, CI/CD, Python, R, SAS
21h
Save
Mark Applied
Hide
Senior Analytics Engineer
Wheat Ridge, Colorado, United States
$94k-$117k/yr RemoteFull Time
Jefferson Center
Jefferson Center: Provides community-based mental health and substance use recovery services.
5+ YOEBachelor's degree or equivalent practical experience, 5+ years in analytics or data engineering, 3+ years in healthcare analytics, and expertise in SQL, Python, R, data modeling, cloud platforms, and AI tools.
SQL, Python, R, Azure Data Factory, Microsoft Fabric, dbt, Snowflake, Databricks, AzureML, REST APIs, GitHub Copilot, Claude, Power BI, Microsoft Fabric Copilots, RAG
23h
Save
Mark Applied
Hide
Analytics Engineer - Predictive Analytics & BI - Ardán Inc.
Maitland, Florida, United States
OnsiteFull Time
Ardán
Ardán: Ardán provides technology-driven solutions to simplify real estate closing processes.
Requires analytics or technical data experience, strong SQL, predictive analytics, data modeling, cloud databases, and Python or R. Strong communication and independent problem-solving skills required; AWS Redshift preferred.
SQL, AWS Redshift, Python, R, Power BI, DAX, Power Query, Microsoft Fabric, OneLake, Data Factory, Lakehouse, Warehouse, Notebooks, ETL, ELT, APIs, Salesforce, CRM
1d
Save
Mark Applied
Hide
Senior Analytics Engineer
Lisbon or Paris or New York City
HybridFull Time
Dashlane
Dashlane: Digital password management and credential security software.
5+ YOERequires 5+ years in analytics engineering or equivalent, expert SQL and dbt skills, B2B SaaS acumen, consultative stakeholder management, AI tooling experience, and fluent English.
SQL, dbt, Python, AWS, Redshift, S3, Lambda, Kinesis, Glue, Omni, Airflow, GitLab, Claude Code
1d
Save
Mark Applied
Hide
Analytics Engineer III, Assurance
Buffalo, New York, United States
$100k-$150k/yr RemoteFull Time
ACV
ACVNASDAQ: ACVA: Provides a digital marketplace for wholesale vehicle auction transactions.
3+ YOEBA/BS in a quantitative or related field, 3+ years in analytics engineering, data engineering, or BI, production dbt experience, expert SQL, project ownership, and Git-based workflows.
dbt, SQL, BigQuery, Git, Omni, Looker, Google Cloud Platform
1d
Save
Mark Applied
Hide
Senior Analytics Engineer
Seattle or Tukwila
$112k-$168k/yr HybridFull Time
ThriftBooks
ThriftBooks: Independent online retailer of used and new books.
Experience building production analytics models with dbt, data warehousing and BI visualization; ability to lead technical projects, optimize Snowflake queries, and collaborate with engineers and business partners.
dbt, Fivetran, Stitch, Snowflake, Mode
1d
Save
Mark Applied
Hide
Sr. Analytics Engineer
United States
$160k-$185k/yr RemoteFull Time
Backblaze
BackblazeNASDAQ: BLZE: Independent cloud storage and data backup platform.
8+ YOERequires 8+ years in analytics or data engineering, advanced SQL, production dbt models, cloud data warehouse experience, semantic-layer expertise, software engineering practices, and strong English communication.
dbt, Snowflake, MetricFlow, Cube, LookML, Git, Salesforce, Stripe, NetSuite, Orb, Tableau, Python, Airflow, Dagster, Fivetran