GlobalLogic
Posted 1w ago

Senior Python Data Engineer IRC301719

GlobalLogic
Romania
RemoteFull Time
Responsibilities
  • building ETL jobs
  • maintaining infrastructure
  • diagnosing data-quality issues
Requirements
  • Requires 6+ years of Python development
  • Production AWS Glue and PySpark, SQL
  • Modern table formats
  • Cloud data lakes
  • Infrastructure as code
  • Automated testing, and CI/CD experience
Technical tools mentioned
PythonAWS GluePySparkApache IcebergDelta LakeHudiSQLPulumiTerraformAWS CDKAmazon S3AWS LambdaAWS Step FunctionsAWS IAMAWS Secrets ManagerAWS Lake FormationDatabricksUnity CatalogGitHub ActionsSonarQubeBlackDuckpytest

Job description

Description

Automotive industries



Requirements

About the role
The project is migrating dealer invoice data (Sell-Out Retail: Workshop and invoices) from on-premise DMS feeds into an AWS-native data platform (Hyperion), against a hard program deadline of 15 December 2026. You’ll join the engineering side of the program, working on the data products that ingest, cleanse, and publish this data — with the Data Product SOR fact layer (currently ~81% delivered, only 25% approved by business testers) as the most critical, highest-priority piece of remaining scope.
This isn’t a generic “Python + Spark” role — you’ll be working inside an existing, opinionated internal framework (mesh_cdk) across several sibling data products that all follow the same raw → trusted → shared (medallion) pattern, so ramp-up is about learning that pattern once and applying it everywhere.




Job responsibilities

What you’ll do
• Build and maintain AWS Glue (PySpark) ETL jobs that cleanse, transform, and publish invoice/workshop data into Apache Iceberg tables.
• Work across the DP SOR / DP Invoice / Master Data pipelines — implementing per-market cleansing rules, data-quality checks, and reconciliation logic against Classic and Compact market feeds.
• Write and maintain Pulumi (Python) infrastructure-as-code for Glue jobs, Step Functions, Lambda triggers, and Lake Formation data-sharing permissions.
• Diagnose and fix root-cause data-quality issues in the pipeline itself (not downstream patching) — this program’s stated philosophy, in contrast to the legacy TekCor4 approach.
• Contribute to GDPR-driven data-retention jobs and cross-account data-sharing setups.
• Participate in the existing CI/CD pipeline: GitHub Actions promotion across dev → int → prod, SonarQube/BlackDuck scan gates, pytest-based test suites.
• Work within a market-by-market delivery cadence (features to business every 2 weeks, feedback loop +1 week) — prioritization is largely driven by an active ~189-item backlog.

Must have
• 6+ years professional Python development.
• Hands-on production experience with AWS Glue and PySpark (or equivalent Spark experience you can translate quickly).
• Solid SQL — you’ll be writing and debugging transformation/reconciliation queries regularly.
• Experience with a modern table format (Apache Iceberg, Delta Lake, or Hudi) and a cloud data lake / lakehouse pattern (raw/trusted/shared or bronze/silver/gold).
• Infrastructure-as-code experience — Pulumi preferred, Terraform/CDK acceptable if you can ramp on Pulumi quickly.
• Comfortable working with AWS core services beyond Glue: S3, Lambda, Step Functions, IAM, Secrets Manager.
• Experience writing automated tests (pytest) and working inside a CI/CD pipeline with mandatory quality gates.
• Ability to work independently inside an existing large, multi-repo codebase with established conventions — this is a stabilization/completion effort, not a greenfield build.

Nice to have
• Experience with AWS Lake Formation (cross-account data sharing / fine-grained permissions).
• Familiarity with Databricks / Unity Catalog (one sibling data product uses Databricks instead of Glue).
• Background in automotive, retail, or invoice/financial data domains.
• Experience with data-quality frameworks or GDPR-driven retention/deletion pipelines.
• Prior work in a data-mesh / domain-oriented data product architecture (vs. a monolithic warehouse).

Why this role is real, not theoretical
This isn’t backlog grooming — there’s an active, prioritized 189-story backlog (37 items specifically tagged DP SOR, the highest count of any component) and a fixed external deadline. You’d be picking up work that’s already scoped and triaged, with existing runbooks, an established test/deploy pipeline, and a clear definition of what “done” looks like per market (quality-gate sign-off, not just code merged).

#LI-TR2 #LI-Remote



What we offer

Empowering Projects: With 500+ clients spanning diverse industries and domains, we provide an exciting opportunity to contribute to groundbreaking projects that leverage cutting-edge technologies. As a team, we engineer digital products that positively impact people’s lives.

Empowering Growth: We foster a culture of continuous learning and professional development. Our dedication is to provide timely and comprehensive assistance for every consultant through our dedicated Learning & Development team, ensuring their continuous growth and success.

DE&I Matters: At GlobalLogic, we deeply value and embrace diversity. We are dedicated to providing equal opportunities for all individuals, fostering an inclusive and empowering work environment.

Career Development: Our corporate culture places a strong emphasis on career development, offering abundant opportunities for growth. Regular interactions with our teams ensure their engagement, motivation, and recognition. We empower our team members to pursue their career goals with confidence and enthusiasm.

Comprehensive Benefits: In addition to equitable compensation, we provide a comprehensive benefits package that prioritizes the overall well-being of our consultants. We genuinely care about their health and strive to create a positive work environment.

Flexible Opportunities: At GlobalLogic, we prioritize work-life balance by offering flexible opportunities tailored to your lifestyle. Explore relocation and rotation options for diverse cultural and professional experiences in different countries with our company.


About GlobalLogic

GlobalLogic, a Hitachi Group Company, is a trusted digital engineering partner to the world’s largest and most forward-thinking companies. Since 2000, we’ve been at the forefront of the digital revolution – helping create some of the most innovative and widely used digital products and experiences. Today we continue to collaborate with clients in transforming businesses and redefining industries through intelligent products, platforms, and services.

About GlobalLogic

Digital product engineering and software development services provider.

Similar jobs

Data Engineer roles
7h
Save
Mark Applied
Hide
Data Engineer - 6 month contract (remote)
Tallinn or Bucharest or Barcelona or Lisbon or Warsaw or Kyiv or Riga
RemoteFull Time, Contract
FYUL
FYUL: A platform powering global on-demand eCommerce merchandise production.
4+ YOERequires 4+ years in data or analytics engineering, or 3+ years with financial data experience; Snowflake or equivalent, advanced SQL, transactional datasets, independent delivery, and strong written communication.
Snowflake, SQL, Looker, LookML, Microsoft Dynamics, dbt, Python
7h
Save
Mark Applied
Hide
Data Engineer P&C
Bucharest or Paris
OnsiteContract
SCOR
SCOREuronext Paris: SCR: Global reinsurance providing risk management and insurance solutions.
3+ YOERequires 3+ years of data engineering experience, data pipeline development, and proficiency in Python, PySpark, and SQL. Databricks, Palantir Foundry, CI/CD, Gitflows, REST APIs, and Power BI are advantageous.
Databricks, Palantir Foundry, Python, PySpark, SQL, Scrum, Kanban, CI/CD, Gitflows, REST API, Power BI
8h
Save
Mark Applied
Hide
Senior Data Engineer (GCP)
Cluj-Napoca, Cluj, Romania
HybridFull Time
Endava
EndavaNYSE: DAVA: Provides software engineering and digital transformation consulting services.
5+ YOERequires 5+ years in data engineering or related fields, advanced SQL, GCP production experience, scalable ETL/ELT design, programming in Python, Scala, or Java, and data governance expertise.
Google Cloud Platform (GCP), SQL, dbt, Python, Scala, Java, Apache Airflow, Matillion, Fivetran, Informatica, CI/CD
9h
Save
Mark Applied
Hide
Senior Data Engineer
Bulgaria or Czech Republic or Hungary or Moldova or Romania or Slovakia or Poland
RemoteFull Time
Xebia
Xebia: Global IT consultancy providing software engineering and cloud services.
Large-scale data engineering experience with streaming, batch pipelines, AWS, Kafka, Iceberg, DynamoDB, Flink, SQL or Python, stakeholder collaboration, technical ownership, fluent English, EU work authorization, and Lead Engineer experience.
AWS, FlinkSQL, Streaming, Flink, SQL, Python, Kafka, Iceberg, DynamoDB, Amazon Web Services (AWS), Claude Code, GitHub Copilot, Cursor, Kubernetes, AI
1d
Save
Mark Applied
Hide
Data Engineer - Bucuresti, Romania
Bucharest, Bucharest, Romania
HybridFull Time, Contract
Societe Generale
Societe GeneraleEuronext Paris: GLE: Global provider of retail, investment, and private banking services.
3+ YOERequires 3–5 years in data engineering, strong Python, SQL, and PySpark, ETL/ELT pipeline experience, data modeling, distributed platforms, governance, monitoring, and cross-functional collaboration.
Python, SQL, PySpark, Palantir Foundry, Databricks, Snowflake, Synapse, Azure, AWS, GCP, CI/CD, DevOps, AI/ML
1d
Save
Mark Applied
Hide
Associate Data Engineer - Sportsbet (12 months), Hybrid
Cluj-Napoca, Cluj, Romania
HybridFull Time, Temporary
Flutter Entertainment
Flutter EntertainmentNYSE: FLUT: Provides online sports betting and iGaming services globally.
Experience with Python, SQL, cloud technologies, AWS managed services, agile development, test-driven development, systems analysis, component-based design, and ETL tools.
Python, SQL, AWS, S3, SNS, SQS, EventBridge, CloudWatch, Lambda, DynamoDB, Spectrum, Redshift, EMR, Databricks, Airflow, NiFi, AWS Glue, Udemy
1d
Save
Mark Applied
Hide
Data Engineer
Timisoara, Timiș County, Romania
lei140k/yr HybridFull Time
Magna International
Magna InternationalNYSE: MGA: Designs and manufactures automotive systems and complete vehicle assemblies.
Experience building scalable data pipelines, warehouses, lakes, and cloud platforms; strong programming in Python, Scala, or Java; SQL/NoSQL, data governance, Git, and distributed systems knowledge.
Python, Scala, Java, SQL, NoSQL, Apache Spark, Databricks, Kafka, Azure, AWS, Google Cloud, Docker, Kubernetes, Terraform, Azure Bicep, AWS CDK, GDPR, Power BI, Tableau, Grafana, Git, GitHub, GitLab, ETL, ELT
2d
Save
Mark Applied
Hide
Data Engineer
Bucharest, Bucharest Municipality, Romania
HybridFull Time
AllCloud
AllCloud: Provides cloud consulting, managed services, and software implementation solutions.
4+ YOERequires 4–6 years in data, architecture, or backend engineering; AWS data services; SQL and Python; cloud data platforms, lakehouses, real-time pipelines, CI/CD, IaC, monitoring, and customer collaboration.
AWS, Databricks, DBT, Snowflake, Airflow, Kafka, Spark, Redshift, S3, AWS Glue, Lambda, SQL, Python, Amazon Bedrock, SageMaker, Kubernetes