Tata Consultancy Services
Posted 2w ago

Data Engineer

Tata Consultancy Services
Irving, Texas, United States
$125k-$140k/yrOnsiteFull Time
Responsibilities
  • building pipelines
  • optimizing storage
  • tuning performance
Requirements
  • Requires a bachelor's in computer science and 6–8 years of experience with data engineering, Spark
  • PySpark, Hive
  • Python, SQL
  • Data pipelines
  • Warehousing, and cloud data lakes
Technical tools mentioned
Apache SparkPySparkHiveCloudera PlatformSpark SQLApache HiveSpark UIApache AirflowPythonHiveQLANSI SQLParquetORCAvroGitJenkinsAnsibleAWS EMRAWS DatabricksHBaseCassandraMongoDB

Job description

Roles & Responsibilities

Job Title: Data Engineer  


Job Description:


We are seeking a highly skilled and motivated Data Engineer to play a pivotal role in designing, building, and optimizing our next-generation scalable data pipelines. This position requires expertise in processing massive datasets using cutting-edge technologies like Apache Spark, PySpark, and Hive within Cloudera Platform. Your primary objective will be to ensure the utmost data reliability, speed, and efficiency, providing a robust foundation for downstream business intelligence and advanced analytics initiatives.


Roles & Responsibilities:

• Data Pipeline Development & Maintenance: Design, build, and maintain highly scalable and efficient ETL/ELT data pipelines utilizing PySpark and Spark SQL , Hive for complex data transformations.

• Data Warehousing & Storage Optimization: Strategically manage data layout, partitioning, and indexing within Apache Hive and various cloud data lake solutions to optimize performance and accessibility.

• Performance Tuning & Optimization: Proactively identify and resolve performance bottlenecks in Spark jobs, leveraging Spark UI for in-depth analysis, effectively managing data skewness, and optimizing memory utilization.

• Diverse Data Integration: Develop robust solutions for ingesting high-volume and diverse datasets from both structured relational databases and unstructured flat files into our data ecosystem.

• Automated Workflow Orchestration: Implement and manage automated data workflows using industry-standard scheduling tools like Apache Airflow or platform-native schedulers, ensuring timely and reliable data delivery.

• Strategic Collaboration: Partner closely with data scientists, business analysts, and cross-functional enterprise teams to translate complex business requirements into technically sound and efficient data solutions.

Qualifications:

• Big Data Frameworks Expertise: Demonstrated high proficiency in Apache Spark architecture, including a deep understanding of drivers, executors, and Directed Acyclic Graphs (DAGs).

• Advanced Programming: Exceptional coding skills in Python and extensive experience with the PySpark API for developing intricate data transformations and processing logic.

• Querying & Schema Management: Strong command of HiveQL and ANSI SQL, coupled with expertise in data partitioning techniques and effective schema definition.

• Optimized Storage Formats: In-depth understanding and practical experience with optimized big data storage file formats such as Parquet, ORC, and Avro.

• Data Warehousing Fundamentals: Solid foundation in Dimensional Data Modeling, including Star and Snowflake schemas, and practical experience with Data Lakes concepts and implementation.

Preferred Qualifications

• CI/CD & DevOps Automation: Experience with Continuous Integration/Continuous Deployment (CI/CD) practices and automation tools like Git, Jenkins, or Ansible.

• Cloud Ecosyste m Development: Experience in development experience utilizing cloud-native big data utilities (e.g., AWS EMR, AWS Databricks) within major cloud platforms.

• NoSQL Database Integration: Exposure to and experience with NoSQL databases such as HBase, Cassandra, or MongoDB.

• Professional Certifications: Relevant professional certifications on Spark or Data Engineer are highly valued





Salary Range: $125,000 to $140,000 per year

About Tata Consultancy Services

Global provider of IT services, consulting, and business solutions.

Similar jobs

Data Engineer roles near Irving, Texas
1d
Save
Mark Applied
Hide
Senior Engineer - Data
Dallas, Texas, United States
OnsiteFull Time
Wingstop
WingstopNASDAQ: WING: A fast-casual restaurant chain specializing in chicken wings.
7+ YOEBachelor's degree or equivalent experience, 7+ years of data engineering, expert SQL and Python, cloud data platform expertise, data modeling, pipelines, orchestration, streaming, CI/CD, and testing.
SQL, Python, Snowflake, Databricks, Amazon Redshift, dbt, Airflow, Dagster, Kafka, Amazon Kinesis, Terraform, Git, CI/CD
1d
Save
Mark Applied
Hide
Data Engineer II - AMZ10414442
Dallas, Texas, United States
$139k-$179k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
1+ YOEMaster's degree in a related field and 1 year of experience, or bachelor's degree and 5 years of progressive experience. Requires ETL/ELT, OLAP, data modeling, SQL, and Oracle experience.
ETL, ELT, SQL, Oracle, OLAP, OBIEE
1d
Save
Mark Applied
Hide
Data Engineer-AWS, Python
Richardson or San Antonio or Texas or United States
OnsiteFull Time
Infosys
InfosysNYSE: INFY: Global provider of digital services, consulting, and technology solutions.
Requires Python, Informatica, ETL, AWS, and Glue; SQL and ServiceNow preferred. Requires a bachelor's degree or equivalent progressive experience and unrestricted US work authorization without sponsorship.
Python, Informatica, ETL, AWS, Glue, SQL, ServiceNow
1d
Save
Mark Applied
Hide
Slalom Flex (Project Based) - Data Engineer (Microsoft Fabric)
Atlanta or Austin or Baltimore or Boston or Charlotte or Chicago or Columbus or Dallas or Hartford or Kansas City or Los Angeles or Miami or New York or Philadelphia or Phoenix or Richmond or San Francisco or Seattle or St. Louis or Washington
$50-$70/hr RemoteContract
Slalom
Slalom: Provides business and technology consulting and software engineering services.
Experience with cloud data integration, ETL/ELT, migration, data modeling, Power BI, Microsoft Fabric, testing, validation, and agile collaboration with architects and stakeholders.
Microsoft Fabric, Data Factory, Lakehouse, Warehouse, Power BI, DAX
1d
Save
Mark Applied
Hide
Data Engineer
Dallas, Texas, United States
$63k-$114k/yr HybridFull Time
Kyndryl
KyndrylNYSE: KD: Manages and modernizes mission-critical IT infrastructure systems.
Expertise in data mining, storage, ETL, data pipelines, relational and NoSQL databases, and cloud tooling; strong analytical, problem-solving, communication, and project-management skills.
Retrieval-Augmented Generation (RAG), Pinecone, Milvus, Weaviate, ETL, Glue, Databricks, Synapse, Dataproc, PostgreSQL, DB2, MongoDB, GitHub, Microsoft Visual Studio, Microsoft, Google, Amazon, Skillsoft
1d
Save
Mark Applied
Hide
Senior Data Engineer
Miramar or Dallas
RemoteContract, Full Time
GroupA
GroupA: N/A
5+ YOEBachelor's degree in a related field and 5+ years of data engineering experience. Requires SQL, Python, Pandas, PySpark, data modeling, cloud platforms, orchestration tools, and data application development.
SQL, Python, Pandas, PySpark, Databricks Unity Catalog, dbt, Streamlit, Hex, Great Expectations, Databricks Delta Sharing, OWL, Git
1d
Save
Mark Applied
Hide
Senior Data Engineer
Edmond or Frisco
$140k-$155k/yr HybridFull Time
American Health Staffing Group
American Health Staffing Group: Provides healthcare staffing and workforce management technology solutions.
5+ YOEBachelor's degree in a related technical field and 5–7 years of data or analytics engineering experience, with advanced SQL, ETL/ELT, dbt, Snowflake, data modeling, and reporting expertise.
SQL, dbt, Snowflake, Fivetran, Python, Microsoft Power BI
1d
Save
Mark Applied
Hide
Senior Data Engineer
Edmond or Frisco
$140k-$155k/yr HybridFull Time
Trio Workforce Solutions
Trio Workforce Solutions: Healthcare workforce management platform providing managed services and software.
5+ YOEBachelor’s degree in a related field and 5–7 years of data engineering or related experience. Requires advanced SQL, ETL/ELT, dbt or similar frameworks, Snowflake or similar platforms, data modeling, and troubleshooting skills.
SQL, dbt, Snowflake, Fivetran, Python, Power BI, AI-assisted development tools, ETL, ELT, DevOps