LTIMindtree
Posted 3w ago

Specialist - Data Engineering

LTIMindtree
Irving, Texas, United States
$84k-$128k/yrOnsiteFull Time
Responsibilities
  • designing data solutions
  • developing pipelines
  • optimizing workloads
Requirements
  • Requires 6+ years of experience
  • AWS certification
  • Expert Python
  • PySpark, SQL
  • Shell scripting
  • AWS data services
  • Apache Iceberg
  • Apache Hive
  • Data pipelines
  • Modeling, and governance
Technical tools mentioned
Apache SparkHadoopPythonSparkSQLAmazon Web Services (AWS)Amazon S3Amazon EC2Amazon EMRAWS GlueAmazon AthenaAWS LambdaAmazon RedshiftAmazon KinesisApache IcebergApache HiveAmazon VPCAmazon EKSKubernetesBoto3SQLAWS CodePipelineGitHub ActionsGitLab CIApache AirflowGitApache KafkaApache FlinkPrestoshell scripting

Job description

Job Details

  • Location: Irving - Texas - USA
  • Experience: 6 - 10 Years
  • Job Type: Permanent
  • Openings: 1
  • Compensation: Compensation range: $ 83,912.00 to 128,080.00 per year

Role Description

Job Title: AWS Pyspark Developer

Work Location : Irving,Texas

Job Summary

 

  • We are seeking a highly skilled and motivated AWS Certified Engineer to design build and optimize scalable data solutions within the Amazon Web Services AWS ecosystem The ideal candidate will have strong expertise in big data processing using PySpark and a deep understanding of data warehousing concepts including Hive and modern table formats like Iceberg This role involves developing deploying and managing robust efficient and secure data pipelines and analytics solutions on AWS leveraging core networking and compute services

 Responsibilities

  • AWS Solution Design Implementation Design develop and deploy scalable and costeffective data solutions on AWS leveraging services such as S3 for data lakes EC2 EMR Glue Athena Lambda Redshift and Kinesis
  • Data Pipeline Development Build and maintain robust ETLELT data pipelines using PySpark for data ingestion transformation and loading into various data stores including those utilizing open table formats like Iceberg
  • Big Data Processing Develop and optimize big data processing jobs using PySpark on AWS EMR or AWS Glue handling large datasets efficiently and integrating with Iceberg table formats
  • Data Warehousing Design implement and manage data warehousing solutions including schema design data modeling and query optimization with a focus on Hive and modern data lake table formats like Iceberg for historical data and analytical queries
  • Cloud Infrastructure Networking Implement secure and robust cloud infrastructure components including VPCs subnets routing and security groups to ensure proper connectivity and isolation for data solutions
  • Containerized Workloads Design deploy and manage containerized data processing applications on Amazon Elastic Kubernetes Service EKS
  • Performance Tuning Optimization Optimize AWS resources and big data applications Spark Hive Iceberg for performance cost and efficiency
  • Data Governance Security Implement best practices for data security access control and compliance within AWS including IAM policies S3 bucket policies and encryption
  • Monitoring Troubleshooting Set up monitoring ing and logging for data pipelines and AWS infrastructure troubleshoot and resolve issues promptly
  • Automation Develop and maintain automation scripts using Python and shell scripting for infrastructure provisioning deployment and operational tasks
  • Collaboration Work closely with data scientists analysts and other engineering teams to understand data requirements and deliver reliable data solutions

Required Skills , Qualifications

  • AWS Certification Hold at least one AWS certification eg AWS Certified Solutions Architect Associate AWS Certified Data Analytics Specialty AWS Certified Developer Associate
  • AWS Services Expertise Handson experience with key AWS services for data processing and storage including
  • Storage S3 for data lakes EC2
  • Data Processing EMR Glue Athena Lambda
  • Networking VPC Subnets Routing Security Groups
  • Containerization EKS
  • Big Data Processing Strong proficiency in PySpark for developing complex data transformations and analytics
  • Data Lake Table Formats Practical experience with Apache Iceberg for managing and querying data lakes
  • Data Warehousing Indepth knowledge and practical experience with Apache Hive for data storage querying and schema management

Programming Languages

  • Python Expertlevel proficiency in Python for scripting data manipulation and AWS automation Boto3
  • Shell Scripting Proficient in shell scripting for automation and operational tasks
  • Database SQL Strong SQL skills for data querying and manipulation
  • Data Concepts Solid understanding of ETLELT processes data modeling distributed computing and data governance
  • Good to Have Skills
  • Containerization Orchestration Experience with Kubernetes for deploying and managing containerized applications
  • CICD Experience with CICD tools and practices eg AWS CodePipeline GitHub Actions GitLab CI for automating deployment of data solutions
  • Orchestration Experience with workflow orchestration tools like Apache Airflow
  • Version Control Proficient in using Git for source code management
  • Other Big Data Technologies Exposure to other big data technologies like Apache Kafka Flink or Presto
  • Certifications
  • AWS Certified Solutions Architect AssociateProfessional
  • AWS Certified Data Analytics Specialty
  • AWS Certified Developer Associate

 

Skills

Mandatory Skills : Apache Spark, Big Data Hadoop Ecosystem, Python, Python for DATA, SparkSQL

About LTIMindtree

Global technology consulting and digital solutions.

Year founded
1996
Employees
90000
Organization type
Public
Headquarters
IN

Similar jobs

Data Engineer roles near Irving, Texas
18h
Save
Mark Applied
Hide
Principal Engineer Data Engineering
Irvine or Alpharetta or Irving or Colorado Springs or Livingston or Tampa or Rolling Meadows
$121k-$231k/yr RemoteFull Time
Verizon
VerizonNYSE: VZ: Provides wireless, broadband, and telecommunications services to customers.
6+ YOEBachelor’s degree or equivalent experience; 6+ years relevant experience, including 5+ years ETL/ELT development; expert SQL and Python; data modeling, warehousing, architecture, and AWS or GCP expertise.
ETL, ELT, Teradata, SQL, Python, AWS, GCP, Apache Airflow, Talend, Git, Jenkins, AI/ML
1d
Save
Mark Applied
Hide
Data Engineer
Plano or Teaneck
$65k/yr OnsiteFull Time
Cognizant
CognizantNASDAQ: CTSH: Provides IT consulting and technology services to global enterprises.
Bachelor's or master's degree in a related field; Python and SQL skills; data orchestration, cloud, data architecture, containerization, CI/CD, analytical, problem-solving, and communication skills.
Python, SQL, Apache Spark, AWS, Microsoft Azure, Google Cloud Platform (GCP), Apache Airflow, Prefect, Snowflake, Databricks, BigQuery, Docker, Kubernetes, GitHub
1d
Save
Mark Applied
Hide
Data Engineer
Plano or Teaneck
$65k/yr OnsiteFull Time
Cognizant
CognizantNasdaq: CTSH: Provides global information technology and business process outsourcing services.
Bachelor's or master's degree in a relevant field; Python and SQL skills; data pipelines, orchestration, cloud, data architecture, containerization, CI/CD, and ETL/ELT knowledge.
Python, SQL, Spark, AWS, Azure, GCP, Airflow, Prefect, Snowflake, Databricks, BigQuery, Docker, Kubernetes, GitHub, JSON, XML, ETL, ELT, CI/CD
1d
Save
Mark Applied
Hide
Senior Data Engineer
Dallas, Texas, United States
HybridContract, Full Time
Parkland Health
Parkland Health: Public academic medical center providing hospital and health services.
3+ YOERequires 3+ years in data engineering, database administration, or related work; ETL, data pipelines, enterprise modeling, SQL, database optimization, cloud environments, and healthcare data experience preferred.
SQL, Microsoft Azure, Microsoft SQL Server, Power BI, Tableau, Python
1d
Save
Mark Applied
Hide
Senior Data Engineer
Dallas, Texas, United States
HybridFull Time, Contract
Parkland Health
Parkland Health: Operates a public hospital system providing comprehensive healthcare services.
3+ YOEBachelor's degree preferred or equivalent experience; 3+ years in data engineering or database administration; ETL, enterprise pipelines, data modeling, database optimization, SQL, cloud, and healthcare data experience.
Microsoft Azure, SQL Server, SQL, Power BI, Tableau, Python
1d
Save
Mark Applied
Hide
Lead Data Engineer
Irving or Miami
OnsiteFull Time
Lennar
LennarNYSE: LEN: Builds and sells quality residential homes across the United States.
8+ YOERequires 8+ years in data engineering, including architecture, modeling, warehousing, ETL, testing, and SDLC practices; 3+ years with AWS, Snowflake, dbt, SQL, and Python; 1+ year with orchestration tools or Qlik Replicate.
Amazon Web Services (AWS), Amazon S3, Amazon EC2, Amazon EMR, Amazon EKS, AWS Glue, AWS Lambda, AWS AppFlow, Amazon CloudWatch, Snowflake Data Cloud, dbt, dbt Cloud, GitHub, SQL, Python, Prefect, Apache Airflow, Qlik Replicate, REST APIs, CI/CD
1d
Save
Mark Applied
Hide
Principal Engineer Data Engineering
Irvine or Colorado Springs or Irving or Livingston or Tampa or Rolling Meadows or Alpharetta
$121k-$231k/yr RemoteFull Time
Verizon
VerizonNYSE: VZ: Global provider of wireless, internet, and communication services.
6+ YOEBachelor's degree or 4+ years' experience; 6+ years relevant experience; 5+ years ETL/ELT development; expert SQL and Python; data modeling, warehousing, architecture, and AWS or GCP expertise.
ETL, ELT, Teradata, DevOps, CI/CD, SQL, Python, AWS, GCP, Apache Airflow, Talend, Git, Jenkins, AI/ML
1d
Save
Mark Applied
Hide
Data Engineer
Mesa or Plano or Teaneck
$65k/yr HybridFull Time
Cognizant
CognizantNasdaq: CTSH: Provides IT consulting and digital business process services.
Bachelor's or master's degree in a related field; Python and SQL skills; familiarity with orchestration, cloud, data platforms, containers, CI/CD, ETL/ELT, data modeling, and architecture.
Python, SQL, Apache Spark, Apache Airflow, Prefect, Snowflake, Databricks, BigQuery, Docker, Kubernetes, GitHub, AWS, Microsoft Azure, Google Cloud Platform