📋 External Recruiting Agencies

Infotree Global is a staffing and recruitment agency, and the job description explicitly mentions it is for a position with one of their customers, not a direct hire for the agency itself.

This company was flagged and excluded from default search results. Proceed with caution.

Infotree Global Solutions
Posted 1w ago

Data Engineer AI Java/Python/Spark

Infotree Global Solutions
Warsaw, Masovian Voivodeship, Poland
HybridFull Time
Responsibilities
  • developing software
  • designing pipelines
  • building platforms
Requirements
  • Requires 5+ years of software or data engineering experience
  • Python or Java
  • Apache Spark
  • Distributed systems
  • Data pipelines
  • Databricks
  • Snowflake
  • Kubernetes
  • GenAI/LLM applications, and strong software engineering fundamentals
Technical tools mentioned
PythonJavaApache SparkDatabricksSnowflakeKubernetesLangChainLangGraphAWSMicrosoft AzureGoogle Cloud PlatformApache KafkaDelta LakeDockerLLMsRAGCI/CDAPIs

Job description

Role Overview:

We are looking for a Senior/Lead Data & GenAI Engineer to design, build and maintain scalable, production-grade data and AI platforms.

The role combines software engineering, distributed data processing, cloud-native technologies and Generative AI. You will work across teams, drive technical initiatives and build reusable libraries and frameworks that enable reliable, scalable and testable systems.

Key Responsibilities:

  • Develop, test and maintain high-quality, production-ready software.

  • Design and implement large-scale data pipelines and distributed processing systems.

  • Build scalable cloud-native services and platforms using modern engineering practices.

  • Provide technical leadership for cross-team initiatives and complex engineering projects.

  • Design and develop reusable libraries, frameworks and platform components.

  • Optimize distributed data processing workloads for performance, scalability and reliability.

  • Work with data platforms including Databricks, Apache Spark and Snowflake.

  • Develop and deploy applications using Python and/or Java.

  • Build and operate containerized workloads using Kubernetes and cloud-native technologies.

  • Design and implement GenAI/LLM-based applications and services.

  • Work with frameworks such as LangChain and LangGraph for LLM orchestration and agentic workflows.

  • Collaborate with data scientists, software engineers, architects and product teams.

  • Establish engineering best practices around testing, observability, reliability and deployment.

Required Experience:

  • 5+ years of professional software/data engineering experience.

  • Strong hands-on experience with Python and/or Java.

  • Strong experience with Apache Spark and distributed data processing.

  • Experience with Databricks and/or modern lakehouse platforms.

  • Experience with Snowflake or comparable cloud data warehouses.

  • Practical experience with Kubernetes and cloud-native technologies.

  • Experience designing and maintaining large-scale data pipelines.

  • Strong understanding of distributed systems, scalability and production engineering.

  • Experience developing ML/AI or GenAI applications.

  • Experience with LLM-based applications, RAG, AI agents or LLM orchestration.

  • Familiarity with LangChain, LangGraph or similar GenAI frameworks.

  • Strong software engineering fundamentals including testing, code quality and system design.

Nice to Have:

  • Experience with AWS, Azure or GCP.

  • Experience with streaming technologies such as Kafka.

  • Experience with Delta Lake / Lakehouse architecture.

  • Experience building RAG pipelines and vector-search solutions.

  • Experience with LLM evaluation, observability and productionization.

  • Experience with AI agents, tool calling and multi-step workflows.

  • Experience building internal developer platforms, frameworks or reusable engineering libraries.

  • Experience leading cross-functional or cross-team technical initiatives.

Ideal Candidate Profile:

The strongest candidate is not purely a Data Engineer and not purely an ML Engineer.

We are looking for someone who combines:

Software Engineering + Data Engineering + Cloud/Platform Engineering + GenAI

Typical backgrounds may include:

  • Senior Data Engineer

  • Lead Data Engineer

  • Senior Software Engineer – Data

  • Data Platform Engineer

  • Senior Cloud Data Engineer

  • AI/ML Platform Engineer

  • Senior ML Engineer with strong data engineering experience

  • GenAI Engineer with strong distributed-data/platform experience

  • Data & AI Architect / Technical Lead

Core Technology Stack:

Languages: Python, Java

Data: Apache Spark, Databricks, Snowflake, Delta Lake

Cloud/Platform: Kubernetes, Docker, AWS/Azure/GCP, cloud-native technologies

GenAI/ML: LLMs, RAG, LangChain, LangGraph, AI agents, vector search

Engineering: Distributed systems, APIs, CI/CD, automated testing, observability, scalability

Similar jobs

Data Engineer roles near Warsaw, Masovian Voivodeship
10h
Save
Mark Applied
Hide
Data Engineer - 6 month contract (remote)
Tallinn or Bucharest or Barcelona or Lisbon or Warsaw or Kyiv or Riga
RemoteFull Time, Contract
FYUL
FYUL: A platform powering global on-demand eCommerce merchandise production.
4+ YOERequires 4+ years in data or analytics engineering, or 3+ years with financial data experience; Snowflake or equivalent, advanced SQL, transactional datasets, independent delivery, and strong written communication.
Snowflake, SQL, Looker, LookML, Microsoft Dynamics, dbt, Python
13h
Save
Mark Applied
Hide
Senior Data Engineer
Warsaw, Masovian Voivodeship, Poland
RemoteFull Time
Sigma Software
Sigma Software: Global IT services provider offering custom software development and consulting.
5+ YOE5+ years in data engineering; strong Python and SQL; Spark/PySpark; Databricks or Snowflake; cloud experience; streaming, ETL/ELT, orchestration, IaC, and upper-intermediate English.
Python, SQL, Apache Spark, PySpark, Databricks, Snowflake, Azure, AWS, GCP, Kafka, Spark Structured Streaming, Apache Airflow, Azure Data Factory, Terraform, Unity Catalog, Apache Atlas, dbt, MLflow
2d
Save
Mark Applied
Hide
Data Engineer (Scala) - Data Assets
Warsaw or Kraków
zł15k-zł21k/mo HybridFull Time
Allegro
AllegroWarsaw Stock Exchange: ALE: Operates an online marketplace for buying and selling consumer products.
Programming experience in Scala, Java, or Python; distributed systems and data processing knowledge; cloud experience; Unix/Linux proficiency; clean code, TDD, CI/CD, teamwork, and B2 English.
Scala, Java, Python, dbt, Spark, Apache Beam, Google Cloud Platform (GCP), Dataproc, GKE, BigQuery, Pub/Sub, Dataflow, Composer, Azure, AWS, Unix, Linux, TDD, CI/CD, Kubernetes, Docker, Consul, GitHub, GitHub Actions, Hermes
2d
Save
Mark Applied
Hide
Data Engineer
Warsaw, Masovian Voivodeship, Poland
zł21k-zł30k/mo HybridFull Time, Contract
Asana
AsanaNYSE: ASAN: Software platform for team project management and workflow automation.
3+ YOE3+ years in data or software engineering; computer science, engineering, or equivalent experience; Databricks, Spark, Airflow, SQL, and a modern programming language such as Python, Scala, or Java.
Databricks, Apache Spark, Apache Airflow, SQL, Python, Scala, Java, MacBook, Modern Health, Carrot
6d
Save
Mark Applied
Hide
Senior Data Engineer in Risk Reporting
Warsaw, Masovian Voivodeship, Poland
zł10k-zł22k/mo OnsiteFull Time
ING
INGEuronext Amsterdam: INGA: Provides retail and wholesale banking services to global customers.
3+ YOEMaster’s degree or equivalent experience; 3–5 years in data reporting; data management, modeling, and reporting solutions; SAS 4GL/Python; Power BI, SAS Visual Analytics, or IBM Cognos; English proficiency.
SAS 4GL, Python, Microsoft Power BI, SAS Visual Analytics, IBM Cognos
6d
Save
Mark Applied
Hide
Senior Data Engineer in Risk Reporting
Warsaw, Masovian Voivodeship, Poland
zł10k-zł22k/mo OnsiteFull Time
ING
INGEuronext Amsterdam: INGA: Provides retail, wholesale, and investment banking services worldwide.
3+ YOEMaster's degree or equivalent experience; 3–5 years in data reporting; data management, modeling, and reporting solutions; SAS 4GL/Python; Power BI, SAS Visual Analytics, or IBM Cognos; English proficiency.
SAS 4GL, Python, Power BI, SAS Visual Analytics, IBM Cognos
6d
Save
Mark Applied
Hide
Software Engineer - Lakehouse and AI Data Platform - Warsaw - Analyst / Associate
Warsaw, Mazowieckie, Poland
OnsiteFull Time
Goldman Sachs
Goldman SachsNYSE: GS: Global investment banking, securities, and investment management firm.
Bachelor's or master's degree or equivalent experience; strong Python or Java, SQL, data modelling, production pipelines, Apache Spark, data quality, distributed processing, and software engineering practices.
Python, Java, SQL, Apache Spark, JSON, Avro, Parquet, Kafka, Snowflake, Apache Iceberg, Databricks, Hadoop, Sybase IQ, CI/CD, Kubernetes
1w
Save
Mark Applied
Hide
Data Engineer
Warsaw, Masovian Voivodeship, Poland
HybridFull Time
Bunge
BungeNew York Stock Exchange: BG: Connecting farmers to consumers with agricultural commodities and food ingredients.
3+ YOEBachelor's degree or equivalent experience; 3–5 years related experience; Agile/Scrum collaboration; SQL, database, Python, PL/SQL, BI, ETL, cloud, and data pipeline expertise.
SQL, GitHub, Oracle, BigQuery, PostgreSQL, Microsoft SQL Server, MySQL, Python, PL/SQL, VSC GitHub Copilot, Tableau, Tableau Pulse, Tableau Next, AecorSoft, ETL