KPMG Global Services
Posted 2w ago

Engineer (A2 DES, Databricks, PySpark, Python)

KPMG Global Services
Bangalore, Karnataka, India
OnsiteFull Time
Responsibilities
  • developing solutions
  • building pipelines
  • optimizing performance
Requirements
  • B.Tech/B.E/MCA in CS/IT,2-4 years data engineering experience with Databricks
  • PySpark
  • Python
  • Spark SQL
  • Hands-on Azure Databricks, ADF, ADLS
  • ETL, data pipelines, and performance tuning
Technical tools mentioned
Azure DatabricksDatabricksPySparkPythonSpark SQLAzure Data Factory (ADF)Azure Data Lake Storage (ADLS)Unity CatalogApache SparkMicrosoft FabricAzure AI servicesGenerative AI

Job description

Roles & responsibilities

Role Overview: The Associate 2 - “Data Engineer with Databricks/Python skills” will be part of the GDC Technology Solutions (GTS) team, working in a technical role in the Audit Data & Analytics domain that requires developing expertise in KPMG proprietary D&A (Data and analytics)) tools and audit methodology. He/she will be a part of the team responsible for extracting and processing datasets from client ERP systems (SAP/Oracle/Microsoft Dynamics) or other sources to provide insights through data warehousing, ETL and dashboarding solutions to Audit/internal teams and be involved in developing solutions using a variety of tools & technologies

The Associate 2 - “Data Engineer” will be predominantly responsible for:

Data Engineering


·Understand requirements, validate assumptions, and develop solutions using Azure Databricks, Azure Data Factory or Python. Able to handle any data mapping changes and customizations within Databricks using PySpark

·Build Azure Databricks notebooks to perform data transformations, create tables, and ensure data quality and consistency. Leverage Unity Catalog for data governance and maintaining a unified data view across the organization

·Analyze enormous volumes of data using Azure Databricks and Apache Spark. Create pipelines and workflows to support data analytics, machine learning, and other data-driven applications

·Able to integrate Azure Databricks with ERP systems or third part systems using APIs and build Python or PySpark notebooks to apply business transformation logic as per the common data model

·Debug, optimize and performance tune and resolve issues, if any, with limited guidance, when processing large data sets and propose possible solutions

·Must have experience in concepts like Partitioning, optimization, and performance tuning for improving the performance of the process

·Implement best practices of Azure Databricks design, development, Testing and documentation

·Work with Audit engagement teams to interpret the results and provide meaningful audit insights from the reports

·Participate in team meetings, brainstorming sessions, and project planning activities

·Stay up-to-date with the latest advancements in Azure Databricks, Cloud and AI development, to drive innovation and maintain a competitive edge

·Enthusiastic to learn and use Azure AI services in business processes.

·Work experience on using Microsoft Fabric is an added advantage

·Write production ready code

·Design, develop, and maintain scalable and efficient data pipelines to process large datasets from various sources using Azure Data Factory (ADF).

·Integrate data from multiple data sources and ensure data consistency, quality, and accuracy, leveraging Azure Data Lake Storage (ADLS).

·Design and implement ETL (Extract, Transform, Load) processes to ensure seamless data flow across systems using Azure

·Work experience on Microsoft Fabric is an added advantage

·Enthusiastic to learn, adapt and integrate Gen AI into the business process and should have experience working with Azure AI services

·Optimize data storage and retrieval processes to enhance system performance and reduce latency.

Technical Skills

Primary Skills:


Ø2-4 years of experience in data engineering, with a strong focus on Databricks, PySpark, Python and Spark SQL.

ØProven experience in implementing ETL processes and data pipelines

ØHands-on experience with Azure Databricks, Azure Data Factory (ADF), Azure Data Lake Storage (ADLS)

ØAbility to write reusable, testable, and efficient code

ØDevelop low-latency, high-availability, and high-performance applications

ØUnderstanding of fundamental design principles behind a scalable application

ØGood knowledge of Azure cloud services

ØFamiliarity with Generative AI and its applications in data engineering

ØKnowledge of Microsoft Fabric and Azure AI services is an added advantage

 Enabling Skills


·Excellent analytical, and problem-solving skills

·Quick learning ability and adaptability

·Effective communication skills

·Attention to detail and good team player

·Willingness and ability to deliver within tight timelines

·Flexible to work timings and willingness to work on different projects/technologies

Responsibilities

Roles & responsibilities

Role Overview: The Associate 2 - “Data Engineer with Databricks/Python skills” will be part of the GDC Technology Solutions (GTS) team, working in a technical role in the Audit Data & Analytics domain that requires developing expertise in KPMG proprietary D&A (Data and analytics)) tools and audit methodology. He/she will be a part of the team responsible for extracting and processing datasets from client ERP systems (SAP/Oracle/Microsoft Dynamics) or other sources to provide insights through data warehousing, ETL and dashboarding solutions to Audit/internal teams and be involved in developing solutions using a variety of tools & technologies

The Associate 2 - “Data Engineer” will be predominantly responsible for:

Data Engineering


·Understand requirements, validate assumptions, and develop solutions using Azure Databricks, Azure Data Factory or Python. Able to handle any data mapping changes and customizations within Databricks using PySpark

·Build Azure Databricks notebooks to perform data transformations, create tables, and ensure data quality and consistency. Leverage Unity Catalog for data governance and maintaining a unified data view across the organization

·Analyze enormous volumes of data using Azure Databricks and Apache Spark. Create pipelines and workflows to support data analytics, machine learning, and other data-driven applications

·Able to integrate Azure Databricks with ERP systems or third part systems using APIs and build Python or PySpark notebooks to apply business transformation logic as per the common data model

·Debug, optimize and performance tune and resolve issues, if any, with limited guidance, when processing large data sets and propose possible solutions

·Must have experience in concepts like Partitioning, optimization, and performance tuning for improving the performance of the process

·Implement best practices of Azure Databricks design, development, Testing and documentation

·Work with Audit engagement teams to interpret the results and provide meaningful audit insights from the reports

·Participate in team meetings, brainstorming sessions, and project planning activities

·Stay up-to-date with the latest advancements in Azure Databricks, Cloud and AI development, to drive innovation and maintain a competitive edge

·Enthusiastic to learn and use Azure AI services in business processes.

·Work experience on using Microsoft Fabric is an added advantage

·Write production ready code

·Design, develop, and maintain scalable and efficient data pipelines to process large datasets from various sources using Azure Data Factory (ADF).

·Integrate data from multiple data sources and ensure data consistency, quality, and accuracy, leveraging Azure Data Lake Storage (ADLS).

·Design and implement ETL (Extract, Transform, Load) processes to ensure seamless data flow across systems using Azure

·Work experience on Microsoft Fabric is an added advantage

·Enthusiastic to learn, adapt and integrate Gen AI into the business process and should have experience working with Azure AI services

·Optimize data storage and retrieval processes to enhance system performance and reduce latency.

Technical Skills

Primary Skills:


Ø2-4 years of experience in data engineering, with a strong focus on Databricks, PySpark, Python and Spark SQL.

ØProven experience in implementing ETL processes and data pipelines

ØHands-on experience with Azure Databricks, Azure Data Factory (ADF), Azure Data Lake Storage (ADLS)

ØAbility to write reusable, testable, and efficient code

ØDevelop low-latency, high-availability, and high-performance applications

ØUnderstanding of fundamental design principles behind a scalable application

ØGood knowledge of Azure cloud services

ØFamiliarity with Generative AI and its applications in data engineering

ØKnowledge of Microsoft Fabric and Azure AI services is an added advantage

 Enabling Skills


·Excellent analytical, and problem-solving skills

·Quick learning ability and adaptability

·Effective communication skills

·Attention to detail and good team player

·Willingness and ability to deliver within tight timelines

·Flexible to work timings and willingness to work on different projects/technologies

Qualifications

Education Requirements


·B. Tech/B.E/MCA (Computer Science / Information Technology)

Primary Skills:


Ø2-4 years of experience in data engineering, with a strong focus on Databricks, PySpark, Python and Spark SQL.

ØProven experience in implementing ETL processes and data pipelines

ØHands-on experience with Azure Databricks, Azure Data Factory (ADF), Azure Data Lake Storage (ADLS)

ØAbility to write reusable, testable, and efficient code

ØDevelop low-latency, high-availability, and high-performance applications

ØUnderstanding of fundamental design principles behind a scalable application

ØGood knowledge of Azure cloud services

ØFamiliarity with Generative AI and its applications in data engineering

ØKnowledge of Microsoft Fabric and Azure AI services is an added advantage

About KPMG Global Services

Provides global tax, audit, and advisory support services.

Similar jobs

Data Engineer roles near Bangalore, Karnataka
6h
Save
Mark Applied
Hide
Data Engineer, YouTube Business Organization
Bengaluru, Karnataka, India
OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
3+ YOEBachelor's degree or equivalent practical experience, 3+ years coding and building data pipelines, dimensional models, data infrastructure, and exploratory queries; SQL and data platform experience required.
Flume, Dataflow, Apache Spark, SQL, Extract, Transform, Load (ETL)
9h
Save
Mark Applied
Hide
Data Engineer-I, SmartCommerce
Bengaluru, Karnataka, India
OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
1+ YOEBachelor's degree, 1+ year of data engineering experience, data modeling, warehousing, ETL pipeline development, query and scripting language experience; AWS and ETL tools preferred.
SQL, PL/SQL, DDL, MDX, HiveQL, SparkSQL, Scala, Python, KornShell, AWS, Redshift, Amazon S3, AWS Glue, EMR, Kinesis, Firehose, AWS Lambda, IAM, Informatica, ODI, SSIS, BODI, Datastage
9h
Save
Mark Applied
Hide
Associate Principal - Data Engineering
Bengaluru, Karnataka, India
OnsiteFull Time
LTIMindtree
LTIMindtreeNational Stock Exchange of India: LTIM: Global technology consulting and digital solutions.
11+ YOERequires 11–15 years of experience designing Microsoft Fabric data engineering solutions, pipelines, real-time streaming, warehousing, governance, security, performance tuning, and technical leadership.
Microsoft Fabric
10h
Save
Mark Applied
Hide
DE&A - Core - Advanced Data Engineering - Advanced Data Engineering (Other)
Bangalore or India
OnsiteFull Time
Zensar
ZensarNational Stock Exchange of India: ZENSARTECH: Global technology firm providing digital transformation and infrastructure services.
4+ YOERequires 4–7 years of experience designing, modeling, developing, building, and deploying data engineering solutions and pipelines on GCP, including BigQuery; strong communication skills preferred AI/ML knowledge.
Google Cloud Platform (GCP), BigQuery, AI/ML, Agentic AI
11h
Save
Mark Applied
Hide
Engineer, Data Engineering
Bangalore, Karnataka, India
OnsiteFull Time
News Corp
News CorpNasdaq: NWSA: Global media, publishing, and digital information services.
2+ YOERequires 2–4 years of data engineering experience, GCP expertise, BigQuery, Dataform, Cloud Composer, Dataproc, Python, SQL, data modeling, CI/CD, and infrastructure-as-code experience.
Google Cloud Platform (GCP), BigQuery, BigLake, PySpark, Dataproc, Python, SQL, Cloud Composer, Airflow, Terraform, CircleCI, GitHub, Dataform, Cloud Storage, Cloud Functions
11h
Save
Mark Applied
Hide
GCP Data Engineer
Bengaluru, Karnataka, India
OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
5+ YOEBachelor of Technology required; 5–8 years of experience designing GCP data pipelines, ETL/ELT workflows, data lakes, warehouses, BigQuery solutions, CI/CD, and data governance.
GCP, BigQuery, CI/CD, DevOps
15h
Save
Mark Applied
Hide
Data Engineer - AI Labs
Bengaluru or Bangalore
₹2500k-₹3582k/yr OnsiteFull Time
IDFC First Bank
IDFC First BankNational Stock Exchange of India / BSE Limited: IDFCFIRSTB / 539437: Provides universal banking and financial services to individuals and businesses.
2+ YOEBachelor’s or master’s degree in computer science, data engineering, information systems, or related field; 2+ years building data pipelines and architectures; experience with big data, cloud platforms, GenAI, and unstructured data.
MapReduce, Hive, HDFS, YARN, HBase, MongoDB, DynamoDB, AWS, GCP, Azure, Git, APIs, Apache Spark
21h
Save
Mark Applied
Hide
Senior Data Engineer
Bengaluru, Karnataka, India
OnsiteFull Time
JLL
JLLNYSE: JLL: Global commercial real estate and investment management services.
5+ YOEBachelor's degree in computer science, data engineering, or related field; 5+ years in data engineering or full-stack development; advanced Python, SQL, PySpark, Spark, Databricks, cloud, testing, pipelines, ETL, and DevOps experience.
Python, SQL, PySpark, Spark, Databricks, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform (GCP), AWS Redshift, Google BigQuery, Snowflake, Web Services Description Language (WSDL), REST, GitHub