Innovaccer
Posted 18h ago

4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)

Innovaccer
Noida, Uttar Pradesh, India
OnsiteFull Time
Responsibilities
  • building pipelines
  • tuning performance
  • instrumenting workflows
Requirements
  • B.E
  • B.Tech., or M.Sc. in computer science or a related technical field
  • 5+ years of data engineering experience
  • SQL, Spark
  • Trino/Presto
  • Iceberg
  • Object storage
  • Orchestration, CI/CD
  • Python, or Java
Technical tools mentioned
Apache SparkApache IcebergTrinoPrestoDelta LakeHudiAmazon S3ParquetApache AirflowPythonJavaSQLSpark SQLCI/CDHL7CCDAAWSAmazon BedrockAWS HealthLake

Job description

Engineering at Innovaccer

With every line of code, we accelerate our customers' success, turning complex challenges into innovative solutions. Collaboratively, we transform each data point we gather into valuable insights for our customers. Join us and be part of a team that's turning dreams of better healthcare into reality, one line of code at a time. Together, we're shaping the future and making a meaningful impact on the world.

About the Role

As a Senior Software Engineer on the Lakehouse team, you will build and operate the data pipelines at the heart of Innovaccer's on-premise platform: Spark ingestion jobs landing raw healthcare data into Apache Iceberg, and Trino SQL transforms building the layered tables thatpower analytics and applications. You will work hands-on across the full pipeline surface, from file validation and quarantine at ingestion to query performance and table health in serving.

A Day in the Life

● Build Spark ingestion jobs that land high-volume raw files into Iceberg tables with schema handling, bad-record quarantine, and idempotent batch replay.
● Develop and operate Trino SQL transform pipelines across data layers: validation and
typing, business-rule transforms, MERGE-based deduplication, and aggregate builds.
● Port existing warehouse SQL workloads to Trino and Spark SQL dialects, and validate results against source outputs.
● Automate Iceberg table maintenance: compaction, snapshot expiry, and orphan-file cleanup as scheduled workflows.
● Tune query and pipeline performance: partitioning strategy, file sizing, statistics, and resource-group behavior.
● Instrument pipelines with data-quality checks, reconciliation reports, and alerting, and participate in per-dataset validation during rollout phases.

Requirements

● B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field.

● 5+ years of data engineering experience building production pipelines at scale.

● Strong SQL skills and hands-on experience with Apache Spark for batch processing.

● Experience with Trino/Presto (or a comparable distributed SQL engine) and open table formats: Iceberg preferred, Delta Lake or Hudi acceptable.

● Working knowledge of S3-compatible object storage and columnar file formats (Parquet).

● Experience with workflow orchestration tools (Airflow or equivalent) and CI/CD for data pipelines.

● Professional development experience with Python and/or Java.

● Healthcare data formats (HL7, CCDA, claims files) and regulated-environment experience are pluses.

Benefits

Here’s What We Offer

  • Generous Leaves: Enjoy generous leave benefits of up to 40 days.
  • Parental Leave: Leverage one of industry's best parental leave policies to spend time with your new addition.
  • Sabbatical: Want to focus on skill development, pursue an academic career, or just take a break? We've got you covered.
  • Health Insurance: We offer comprehensive health insurance to support you and your family, covering medical expenses related to illness, disease, or injury. Extending support to the family members who matter most.
  • Care Program: Whether it’s a celebration or a time of need, we’ve got you covered with care vouchers to mark major life events. Through our Care Vouchers program, employees receive thoughtful gestures for significant personal milestones and moments of need.
  • Financial Assistance: Life happens, and when it does, we’re here to help. Our financial assistance policy offers support through salary advances and personal loans for genuine personal needs, ensuring help is there when you need it most.

Innovaccer is an equal-opportunity employer. We celebrate diversity, and we are committed to fostering an inclusive and diverse workplace where all employees, regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, marital status, or veteran status, feel valued and empowered.

Disclaimer: Innovaccer does not charge fees or require payment from individuals or agencies for securing employment with us. We do not guarantee job spots or engage in any financial transactions related to employment. If you encounter any posts or requests asking for payment or personal information, we strongly advise you to report them immediately to our HR department at [email protected]. Additionally, please exercise caution and verify the authenticity of any requests before disclosing personal and confidential information, including bank account details.

About Innovaccer

Innovaccer builds software that helps hospitals, clinics, and health insurance companies make sense of all the scattered data they deal with every day — patient records, appointments, insurance claims, lab results — and turns it into something useful and actionable.

Think of a typical hospital: patient information often lives in a dozen different systems that don't talk to each other well. This causes delays in scheduling, slows down doctors, and makes it harder to catch health problems early. Innovaccer's platform connects all of that data together, and increasingly, uses AI "agents" to actually do some of the manual work for healthcare teams — things like scheduling appointments, preparing insurance paperwork, or drafting clinical notes — so that staff can spend more time with patients and less time on admin work.

Some of the largest healthcare systems in the US — including CommonSpirit Health, Atlantic Health, and Banner Health — use Innovaccer's platform today.

The company is also investing heavily in this direction: in June 2026, Innovaccer signed a multi-year partnership with AWS to run its AI agents at a much larger scale, using AWS's cloud AI tools (Amazon Bedrock) and healthcare-specific data infrastructure (AWS HealthLake). This is a strong signal of where the platform — and this engineering team — is headed next. For more information, visit www.innovaccer.com and check us out on YouTube, Glassdoor, LinkedIn, Instagram, and the Web.

About Innovaccer

Cloud platform for healthcare data activation and clinical analytics.

Year founded
2014
Employees
1500
Organization type
Private
Latest investment
Raised $275.00M Series F (2025) — led by B Capital Group, Banner Health, Kaiser Permanente
Headquarters
US

Similar jobs

Data Engineer roles near Noida, Uttar Pradesh
2h
Save
Mark Applied
Hide
Data Engineer-Data Platforms-Google
Gurgaon, Haryana, India
HybridFull Time
IBM
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
Bachelor's degree required; master's preferred. Experience with Google Cloud data platforms, batch and real-time pipelines, data migration, data layer design, and tools including Airflow, dbt, Spark, Python, and Scala.
Google DataProc, Google DataFlow, Google PubSub, Google BigQuery, Google BigTable, Google Cloud Spanner, Google CloudSQL, Google AlloyDB, Google Cloud Storage, Apache Airflow, dbt, Spark, Python, Scala, Hadoop, Apache Beam, Google Cloud Scheduler, Cloud Composer
11h
Save
Mark Applied
Hide
Lead Data Engineer - Azure + Fabric
Gurugram, Haryana, India
HybridFull Time
EXL
EXLNASDAQ: EXLS: Provides data analytics and digital operations solutions to businesses.
8+ YOERequires 8+ years of data engineering experience, Azure or Microsoft Fabric expertise, Spark/PySpark, Python, SQL, Delta Lake, ETL/ELT, data quality, and a bachelor's degree in a relevant field.
Microsoft Azure, Microsoft Fabric, Microsoft OneLake, Microsoft Fabric Lakehouse, Microsoft Fabric Warehouse, Microsoft Fabric Data Pipelines, Notebooks, Spark, PySpark, Python, SQL, Delta Lake, Power BI, Microsoft Purview, Microsoft Azure DevOps, Git, GCP, CSV, Excel, Microsoft SharePoint, Synapse Pipelines, CI/CD, RBAC
18h
Save
Mark Applied
Hide
Snr Data Engineer
Noida, Uttar Pradesh, India
OnsiteFull Time
Alight
AlightNYSE: ALIT: Provides cloud-based HR, payroll, and benefits administration services.
4+ YOERequires 4–8 years of ETL experience, AWS and Hadoop expertise, Scala or Python, PySpark, Spark, SQL, orchestration tools, data warehousing, CI/CD, GitHub, and strong analytical and communication skills.
Apache Spark, Cloudera, AWS Step Functions, AWS Glue, AWS Lambda, Amazon S3, Amazon Redshift, Hadoop, HDFS, Apache Hive, Apache Kafka, Amazon EMR, Kinesis, PySpark, Spark SQL, Parquet, ORC, Apache Airflow, Control-M, Scala, HiveQL, Impala, SQL, Python, Shell, CI/CD, GitHub, Docker, Amazon ECS, Amazon EKS
18h
Save
Mark Applied
Hide
Senior Data Engineer
Gurugram, Haryana, India
OnsiteFull Time
Iris Software
Iris Software: Provides software engineering and IT consulting services to enterprises.
7+ YOERequires 7–8 years of data engineering experience and expertise in PySpark, Python, SQL, Amazon Kinesis, Databricks Workflows, Delta Lake, data quality, and validation.
Amazon Kinesis, PySpark, Python, SQL, Databricks Workflows, Delta Lake, Databricks, Snowflake, Apache Kafka, Apache Airflow
18h
Save
Mark Applied
Hide
Lead Data Engineer - Data Engineering 4C
Gurgaon, Haryana, India
HybridFull Time
Genpact
GenpactNYSE: G: Provides business process management and digital transformation services.
Bachelor's or master's degree in a related field, data engineering expertise, Big Data and cloud computing skills, technical leadership experience, and proficiency in English.
Databricks Platform, Snowflake, Microsoft Azure, Oracle Database, Apache Spark
3d
Save
Mark Applied
Hide
Data Engineer-Senior II
Bengaluru or Gurugram or Mumbai
OnsiteFull Time
FedEx
FedExNYSE: FDX: Global provider of courier, logistics, and transportation services.
4+ YOEBachelor's degree in computer science, engineering, mathematics, statistics, or similar; 4–7 years' experience; Python, PySpark, SAS, SQL, data modeling, cloud, ETL, and distributed data technologies.
Python, PySpark, SAS, Hadoop, Hive, Spark, Azure, Azure Data Factory, Azure Data Lake Storage, Azure DevOps, Databricks, Delta Lake, Docker, Kubernetes, Terraform, Octopus, SQL, Ab Initio, Informatica, DataStage, Power BI
3d
Save
Mark Applied
Hide
Data Engineer-Senior II
Bengaluru or Mumbai or Gurugram
OnsiteFull Time
FedEx
FedExNYSE: FDX: Global provider of express delivery and logistics services.
4+ YOEBachelor’s degree in a relevant discipline and 4–7 years’ experience. Requires Python, PySpark, SAS, SQL, data pipelines, ETL, cloud platforms, Spark, and data modeling skills.
Python, PySpark, SAS, Hadoop, Hive, Spark, Azure, Azure Data Factory, ADLS Storage, Azure DevOps, Databricks, Delta Lake, Docker, CI/CD, Kubernetes, Terraform, Octopus, SQL, Ab Initio, Informatica, DataStage, Power BI
3d
Save
Mark Applied
Hide
Senior Software Engineer
Noida, Uttar Pradesh, India
OnsiteFull Time
Mastek
MastekNational Stock Exchange of India: MASTEK: Provides digital engineering and cloud transformation services to global enterprises.
3+ YOERequires 3–5 years of data engineering experience, excellent SQL, cloud experience with AWS, GCP, or Azure, ETL/ELT expertise, debugging, and a bachelor's degree in computer science, engineering, or a related field.
SQL, Python, Microsoft Power BI, Tableau, AWS, Google Cloud Platform (GCP), Microsoft Azure