This job has expired

This job posting is no longer active and is not accepting applications. Explore similar roles below!

The University of Texas MD Anderson Cancer Center
Posted 1w ago

Data Engineer - Enterprise Data Engineering & Analytics

The University of Texas MD Anderson Cancer Center
Houston, Texas, United States
$107k-$160k/yrRemoteFull Time
Responsibilities
  • building pipelines
  • curating data
  • training users
Requirements
  • Bachelor's degree required
  • 2 years of relevant healthcare or business experience, and experience with data pipelines
  • Python or Spark
  • Cloud data management, LLMs, and analytics delivery
Technical tools mentioned
FoundryFabricPythonSparkLarge Language Models (LLMs)Epic CogitoClarityCaboodle

Job description

The Data Engineer role is a pivotal position within the Enterprise Data Engineering & Analytics Department, supporting the design, build, and operationalization of integrated data pipelines and analytics solutions that enable MD Anderson’s digital business initiatives. The Data Engineer works across the Context Engine framework to deliver end-to-end data engineering solutions while partnering closely with Enterprise Data Engineering & Analytics teams and other institutional stakeholders.
The Data Engineer contributes to the mission of MD Anderson Cancer Center, a leading institution focused on cancer care, research, education, and prevention. In this role, the Data Engineer helps advance enterprise analytics capabilities by ensuring secure, governed, and reusable data assets that accelerate insights and improve time-to-solution across MD Anderson.
Ideal Candidate Statement
The ideal candidate for the Data Engineer role brings a bachelor’s degree in computer science, preferred advanced education in analytics or computer science, hands-on experience building data pipelines in healthcare or research environments, and at least 2 years of Cloud Data Management Framework (Foundry/Fabric) experience, hands-on use of Large Language Models (LLMs) in real-world projects, Python or Spark development, and analytics delivery. Epic certification and the ability to collaborate across technical and clinical teams are strongly preferred.
Position Information
Salary range based on a 40-hour work week: Minimum $106,500 – Midpoint $133,000 – Maximum $159,500
Work location: Houston, Texas or surrounding area preferred
 

This Data Engineer role offers the opportunity to contribute directly to MD Anderson’s mission by enabling high-quality, governed data that supports clinical, research, and operational analytics across the institution. The position provides exposure to enterprise-scale data engineering initiatives, collaboration with experienced engineering and data science professionals, and opportunities for continued learning and career growth while supporting a balanced and sustainable work environment.
• Employer-paid medical coverage starting day one for employees working 30+ hours/week, plus optional group dental, vision, life, AD&D, and disability insurance.
• Accruals for PTO and Extended Illness Bank, plus paid holidays, wellness, childcare, and other leave options.
• Tuition Assistance Program after six months of service and access to extensive wellness, fitness, and employee resource groups.
• Defined-benefit pension through the Teachers Retirement System, voluntary retirement plans, and employer-paid life and reduced salary protection programs.
Responsibilities
Data Engineering – End-to-End Solution Delivery
• Participate in end-to-end solution delivery that increases information capabilities and realizes data value across the institution
• Build and test end-to-end data pipelines across ingestion, curation, transformation, modeling, and consumption within the Context Engine framework
• Integrate data governance processes across data provenance, security, data quality, ontology, and metadata management
• Participate in planning, architecture, analysis, design, and build of data pipelines in partnership with IS, Data Offices, and Data Governance teams
• Contribute to existing data pipelines spanning acquisition, integration, and consumption for defined use cases
Data Curation, Modeling, and Governance
• Build data curation pipelines including profiling, specification creation, cleansing, transforming, standardizing, mastering, harmonizing, validating, and aggregating data
• Monitor and support data quality across the Context Engine
• Incorporate repeatable solution designs and data models to support reuse and scalability
• Promote effective data management practices and understanding of analytics across the enterprise
Standards, Testing, and System Maintenance
• Adhere to IS division standard operating procedures and all MD Anderson policies
• Maintain build standards and governance oversight sign-off aligned with institutional data strategy
• Participate in documentation preparation for enhancements or new technology
• Perform quality control, testing, and peer review of analytics builds
• Support system updates, releases, change control processes, and after-hours support as required
Education, Training, and Collaboration
• Train data scientists, analysts, end users, and data consumers on data pipelining and preparation techniques
• Assist in establishing training plans and curricula for Context Engine tools
• Provide institutional, department, and one-on-one training on EDEA deliverables
• Support liaison relationships with customers and OneIS partners to deliver effective technical solutions
Innovation and Continuous Improvement
• Explore and promote modern tools, techniques, and architectures to automate data preparation and integration tasks
• Improve productivity by reducing manual and error-prone processes
• Model OneIS values through integrity, partnership, quality, and continuous improvement










EDUCATION

  • Required: Bachelor's Degree
  • Preferred: Bachelor's in computer science, Master's degree Business Analytics, Computer Science, Information Technology, Data Science, or related.

WORK EXPERIENCE

  • Required: 2 years Clinical, relevant healthcare information technology, or relevant business experience. or
  • Required: With preferred degree, no experience required.
  • May substitute required education with years of related experience on a one to one basis.
  • Preferred: 3-5 years creating data pipelines in a healthcare research environment, experience building and maintaining analytical reports and dashboards, problem solving skills and ability to translate business/clinical requirements into reliable data models, analytics & reporting- cloud data management solutions like Foundry, Fabric etc, data pipeline & ETL development -hands on experience designing, building and maintaining pipelines using python/spark, hands-on use of Large Language Models (LLMs) in real-world projects, such as integrating generative AI solutions into applications, workflows, or analytics platforms. Candidates should be familiar with prompt engineering, model evaluation, and responsible AI practices. Experience collaborating with cross-functional teams to deploy and scale LLM-powered features is highly desirable.
  • Preferred certifications: EPIC Cogito, Clarity, Caboodle, Clinical Data Model, etc.

Work location: Prefer a candidate in Houston Texas or surrounding area.

LICENSES AND CERTIFICATIONS

  • Required: EPIC - EPIC Certification Must obtain at least one Epic Data Model certification (Clinical, Access, or Revenue) issued by Epic. within 180 Days

OTHER REQUIREMENTS: Must pass pre-employment skills test as required and administered by Human Resources. 

The University of Texas MD Anderson Cancer Center offers excellent benefits, including medical, dental, paid time offretirement, tuition benefits, educational opportunities, and individual and team recognition.

This position may be responsible for maintaining the security and integrity of critical infrastructure, as defined in Section 113.001(2) of the Texas Business and Commerce Code and therefore may require routine reviews and screening. The ability to satisfy and maintain all requirements necessary to ensure the continued security and integrity of such infrastructure is a condition of hire and continued employment.

It is the policy of The University of Texas MD Anderson Cancer Center to provide equal employment opportunity without regard to race, color, religion, age, national origin, sex, gender, sexual orientation, gender identity/expression, disability, protected veteran status, genetic information, or any other basis protected by institutional policy or by federal, state, or local laws unless such distinction is required by law.http://www.mdanderson.org/about-us/legal-and-policy/legal-statements/eeo-affirmative-action.html

















Additional Information

















  • Requisition ID: 182426








  • Employment Status: Full-Time



























  • Employee Status: Regular










  • Work Week: Days

























  • Minimum Salary:

    US Dollar (USD)

    106,500





















  • Midpoint Salary:

    US Dollar (USD)

    133,000















  • Maximum Salary :

    US Dollar (USD)

    159,500














  • FLSA: exempt and not eligible for overtime pay


















  • Fund Type: Hard


















  • Work Location: Remote (within Texas only)


















  • Pivotal Position: Yes


















  • Referral Bonus Available?: Yes


















  • Relocation Assistance Available?: Yes





































(While navigating through the site, please be sure to disable your pop-up blocker.)


About The University of Texas MD Anderson Cancer Center

Comprehensive cancer treatment, medical research, and education center.

Year founded
1941
Employees
27000
Organization type
Government
Headquarters
US

Similar jobs

Data Engineer roles near Houston, Texas
23h
Save
Mark Applied
Hide
Senior Data Engineer
Salt Lake City or Marietta or Carol Stream or Phoenix or Cypress or DFW Airport
$118k/yr OnsiteFull Time
R.S. Hughes
R.S. Hughes: Distributes industrial supplies and provides custom material converting services.
3+ YOERequires 3+ years in data or analytics engineering, bachelor's degree in computer science or computer engineering, Azure Synapse and Azure SQL experience, SQL, Python or PySpark, ETL/ELT, dimensional modeling, and Power BI.
Azure Synapse Analytics, Azure Logic Apps, Microsoft Graph, Python, PySpark, SQL, Azure SQL Database, Power BI, Microsoft SQL Server
2d
Save
Mark Applied
Hide
Enterprise Data Engineer
Ridgeland or Alabama or Houston or Memphis or Florida or Atlanta
RemoteFull Time
Trustmark
TrustmarkNASDAQ: TRMK: Provides retail and commercial banking, wealth, and insurance services.
4+ YOEBachelor's degree in data or computer science or equivalent certification, 4 years with modern ETL platforms, database, data warehousing, modeling, SQL, Python, and advanced analytical skills.
IBM Datastage, Informatica, Snowpipe, Azure, AWS, SQL, Python
3d
Save
Mark Applied
Hide
Google Senior Data Engineer
Albany or Arlington or Atlanta or Austin or Beaverton or Bentonville or Boston or Carmel or Charlotte or Chicago or Cincinnati or Cleveland or Columbus or Culver City or Denver or Des Moines or Detroit or Hartford or Houston or Irving or Kirkland or Miami or Milwaukee or Minneapolis or Morristown or Mountain View or Nashville or New York City or Oklahoma City or Overland Park or Philadelphia or Pittsburgh or Raleigh or Redmond or Sacramento or San Diego or San Francisco or Scottsdale or Seattle or St. Louis or St. Petersburg or Walnut Creek or United States
$80k-$266k/yr HybridFull Time
Accenture
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
5+ YOERequires 5+ years in data engineering, analytics, or ML; 4+ years with GCP; 5+ years with SQL and pipelines; 3+ years with Python or AI tools; and a bachelor's degree or equivalent.
Google Cloud Platform (GCP), BigQuery, Looker, Vertex AI, Gemini Foundation Models, Gemini Enterprise, Dataflow, Dataproc, Pub/Sub, Cloud Storage, Looker Studio, Model APIs, Embeddings, Dataplex, IAM, SQL, Python, Git
3d
Save
Mark Applied
Hide
Public Health Data Engineer
San Antonio or Houston or Atlanta or United States
$70k-$116k/yr HybridFull Time
Guidehouse
Guidehouse: Provides management and technology consulting services to diverse organizations.
Bachelor's degree in a technical field, foundational data engineering or software experience, programming knowledge, database concepts, cloud familiarity, and ability to obtain and maintain Public Trust.
Python, SQL, PySpark, AWS, Azure, Git, Databricks, Spark, Power BI, Tableau, Kibana, AWS S3, AWS Lambda, AWS Redshift, Azure Data Factory, Azure Functions, Azure Cosmos DB
3d
Save
Mark Applied
Hide
Data Engineer 1
The Woodlands, Texas, United States
OnsiteFull Time
Accelerated Mobile Power
Accelerated Mobile Power: Provides mobile power solutions using gas turbines and generators.
3+ YOEBachelor's degree in MIS, computer science, engineering, IT, or related field, or equivalent experience; 3–5+ years in a data-focused role; Databricks, PySpark, SparkSQL, SQL, and visualization experience.
Databricks, Delta Lake, Delta Live Tables (DLT), Databricks Jobs, Auto Loader, Change Data Capture (CDC), Spark, PySpark, SparkSQL, SQL, Microsoft SQL Server, Power BI, Tableau, Spotfire, Sigma
3d
Save
Mark Applied
Hide
Data Engineer 1
The Woodlands, Texas, United States
OnsiteFull Time
Beusa Energy
Beusa Energy: Provides oil exploration, fracturing, and power generation services.
3+ YOEBachelor's degree in MIS, computer science, engineering, IT, or related field, or equivalent experience; 3-5+ years in a data-focused role; strong Databricks, PySpark, SparkSQL, SQL, and visualization skills.
Databricks, Delta Lake, Delta Live Tables (DLT), Databricks Jobs, Autoloader, CDC, Spark, PySpark, SparkSQL, SQL, Microsoft SQL Server, Power BI, Tableau, Spotfire, Sigma, Z-Ordering, Auto Optimize, OPTIMIZE, VACUUM, MERGE, Change Data Feed, APPLY CHANGES INTO
4d
Save
Mark Applied
Hide
Data Engineer
Houston, Texas, United States
OnsiteFull Time
Foxconn Assembly
Foxconn AssemblyTaiwan Stock Exchange: 2317: Assembles electronics and computer hardware for global technology brands.
3+ YOEBachelor's degree in a related field, 3+ years in data engineering, ETL and data warehousing experience, SQL and Python proficiency, Big Data platform expertise, and native Mandarin proficiency.
Apache, Cicada, Tableau, SQL, Python, Hadoop, Kudu, Hive, Kafka, Spark, Flink, Linux, CentOS, Red Hat, Ubuntu, GitLab, JIRA, HTML5, CSS3, JavaScript, Chrome
4d
Save
Mark Applied
Hide
Senior Data Engineer
Houston, Texas, United States
OnsiteFull Time
WhiteWater Express Car Wash
WhiteWater Express Car Wash: Operates a chain of express exterior car wash locations.
8+ YOEBachelor's degree in computer science, software engineering, or related field, or 8+ years of software engineering experience; strong ETL, database, cloud, programming, API, CI/CD, and collaboration skills.
Python, SQL, DBT, Snowflake, MySQL, SQL Server, Airflow, Git, CI/CD, AWS, Google Cloud, Azure, DigitalOcean, PHP, Node.js, JavaScript, C#, Laravel, Express.js, WordPress, Agile
This job has expired