Mirai
Posted 1mo ago

Senior Data Engineer

Mirai
Riyadh, Riyadh Province, Saudi Arabia
OnsiteFull Time
Responsibilities
  • building pipelines
  • building retrievals
  • modeling datasets
Requirements
  • 8+ years in data engineering
  • Strong SQL and Python (PySpark)
  • Production AWS data stack experience (S3, Glue
  • Athena
  • Redshift)
  • ELT tools
  • Streaming and event-driven systems
  • Vector stores/embedding workflows, and orchestration/IaC knowledge for AI/ML data pipelines
Technical tools mentioned
SQLPythonPySparkS3GlueAthenaRedshiftAirbyteFivetranMeltanoSQSSNSKinesisAmazon MSKDebeziumCubedbt Semantic LayerpgvectorAmazon OpenSearchPineconeWeaviateMilvusParquetApache IcebergDelta LakeHudiAmazon MWAAStep FunctionsDagsterPrefect

Job description

Our Generative AI products are only as good as the data behind them. This role owns that data layer from end to end: the pipelines that bring data in, the transformations that shape it, and the way it reaches retrieval systems, agents, and analytics. The work runs on AWS, and the aim is a single governed source that every consumer can rely on.

We want someone who has already built data pipelines for AI systems, not only for reporting. Preparing data for an LLM or an agent brings its own work around chunking, embeddings, indexing, and keeping content current, and you have done it before. The team is small and spans several languages, so you will own your pipelines and help set the standards the rest of us follow.

WHAT YOU WILL DO

  • Build and run the batch and streaming pipelines that move data from source systems into the lake and through to the warehouse, owning the layers in between from raw to curated, along with their schema, quality, and lineage.
  • Build the data layer behind retrieval: source connectors, document parsing, chunking, embedding generation, and vector indexing, including re-embedding when content changes.
  • Model curated, query-ready datasets and metrics so AI and analytics consumers work from one definition instead of each rebuilding the logic.
  • Add quality checks, validation, and monitoring so problems surface before they reach a model or a user.
  • Apply access control where it belongs: row and column level rules, PII handling, and entitlement-aware datasets, enforced as close to query time as the stack allows.
  • Work with the platform and DevOps engineers to expose data and retrieval as documented, dependable services.
  • Keep storage, compute, and query costs in check, with particular attention to the cost of embedding and vector workloads
  • Review code, write the documentation, and help shape how the team builds its data layer.

Requirements

  • Eight or more years in data engineering overall. That includes hands-on work building data for AI or ML systems such as retrieval, embeddings, or feature data, which can be a more recent part of your background.
  • Strong SQL and strong Python, including PySpark or similar distributed processing.
  • Production experience across the AWS data stack: S3 for the lake, Glue for ETL and the Data Catalog, Athena for serverless query, and Redshift as the warehouse.
  • Hands-on experience with a layered data architecture, whether you call it medallion (bronze, silver, gold), a data lake feeding a warehouse, or a lakehouse, including building the transformation stages that move data from raw to curated.
  • Experience with an ELT or integration tool such as Airbyte, Fivetran, or Meltano, including building or maintaining connectors.
  • Experience with event-driven pipelines using SQS and SNS, and with at least one streaming or change-data-capture technology such as Kinesis, Amazon MSK, or Debezium.
  • Hands-on experience with a semantic or metrics layer over the warehouse, such as Cube or the dbt Semantic Layer.
  • Hands-on experience with at least one vector store and embedding workflow: pgvector, Amazon OpenSearch, Pinecone, Weaviate, or Milvus.
  • Comfort with columnar and open table formats: Parquet together with Apache Iceberg, Delta Lake, or Hudi.
  • Working knowledge of an orchestrator such as Amazon MWAA, Step Functions, Dagster, or Prefect, and enough infrastructure as code to work closely with DevOps.



About Mirai

Develops and produces video games in Saudi Arabia.

Year founded
2024
Employees
60
Industries
Organization type
Private
Latest investment
Corporate Round (2024) — led by Scopely, Savvy Games Group
Headquarters
SA

Similar jobs

Data Engineer roles near Riyadh, Riyadh Province
1d
Save
Mark Applied
Hide
Senior Consultant/Assistant Manager - Tech Consulting - Data & AI - Data Engineer - Riyadh
Riyadh, Riyadh Region, Saudi Arabia
OnsiteFull Time
EY
EY: Global firm providing audit, tax, and professional consulting services.
3+ YOE3–7+ years in data engineering or Data & AI consulting; strong SQL and Python, data modeling, ETL/ELT, orchestration, cloud platforms, governance, and client consulting experience. Bachelor's or Master's degree required.
SQL, Python, Azure, AWS, GCP, Dataiku
1w
Save
Mark Applied
Hide
Senior Specialist - Data Engineering
Riyadh, Riyadh Province, Saudi Arabia
OnsiteFull Time
Qiddiya Investment Company
Qiddiya Investment Company: Developing a massive entertainment, sports, and arts destination.
Degree in Computer Science or similar, prior data engineering experience, proficiency in Python/Java/Scala, SQL/NoSQL, cloud platforms (GCP/AWS/Azure) and data warehousing; cloud data engineering certification is a plus.
Apache Spark, Apache Kafka, Apache Airflow, DataFlow (Apache Beam), DataProc, Data Fusion, Cloud Composer, Python, Java, Scala, SQL, NoSQL, Google BigQuery, Amazon Redshift, AWS S3, Azure Data Lake Storage, MongoDB, Cassandra, AWS Glue, Azure Data Factory
1mo
Save
Mark Applied
Hide
Senior Data Engineer
Riyadh, Riyadh, Saudi Arabia
OnsiteFull Time
73 Strings
73 Strings: AI-powered valuation and monitoring software for private capital.
8+ YOE8+ years in data architecture/engineering with expertise in Snowflake/Databricks, ETL/ELT, REST/GraphQL, streaming platforms, cloud (AWS/Azure/GCP), dbt/Airflow, and client-facing technical leadership.
Snowflake, Databricks, REST, GraphQL, Kafka, Flink, Spark, dbt, Apache Airflow, AWS, Azure, GCP
1mo
Save
Mark Applied
Hide
4261- Data Engineer
Riyadh, Riyadh Province, Saudi Arabia
OnsiteFull Time
Innovaccer
Innovaccer: Cloud platform for healthcare data activation and clinical analytics.
3+ YOE3+ years data engineering experience with Python and SQL, cloud data platform experience (AWS/Azure/GCP), ETL/ELT pipeline design, data modeling (Snowflake/Redshift/Synapse), Spark/Kafka, orchestration (Airflow/Prefect), and familiarity with healthcare standards (HL7, FHIR).
Snowflake, AWS Redshift, Microsoft Azure Synapse, HL7, FHIR, ICD-10, SNOMED CT, AWS, Azure, GCP, Python, SQL, Apache Spark, Apache Kafka, Apache Airflow, Prefect, dbt, AWS Kinesis, Monte Carlo, Great Expectations
2mo
Save
Mark Applied
Hide
Data Engineer
Riyadh, Riyadh Province, Saudi Arabia
OnsiteFull Time
Consertus
Consertus: Global capital program management and infrastructure advisory firm.
5+ YOEBachelor's degree, 5+ years data engineering experience, advanced SQL and Python, ETL/data warehousing and data modeling expertise, experience with cloud/ETL tools, APIs and semi-structured data, strong problem-solving and communication.
SQL, Python, SSIS, Azure Data Factory, Azure, AWS, GCP, APIs, JSON, XML, Airflow, Prefect, Git, Power BI, Tableau, REST, SOAP
4mo
Save
Mark Applied
Hide
Data Engineer
Riyadh, Riyadh, Saudi Arabia
OnsiteMultiple Commitments Available
Devoteam
Devoteam: Global IT consulting firm providing digital transformation and cloud services.
5+ YOE5+ years in Data Quality, Data Governance, or Data Engineering with SQL and scripting (Python, Shell, or Scala); experience with enterprise DQ tools; knowledge of data governance and MD M frameworks; CDMP/IDQ/IDMC certificates preferred.
SQL, Python, Shell, Scala, Informatica DQ/IDMC, Talend DQ, IDQ, MDM, DWH, Data Lake
6mo
Save
Mark Applied
Hide
Data Engineer
Riyadh, Riyadh, Saudi Arabia
OnsiteFull Time
Artefact
Artefact: Provides data engineering, AI implementation, and digital marketing services.
2+ YOE2-5 years experience; Python and SQL; ETL pipelines; cloud services (Azure, AWS, GCP); Spark & Kafka; strong communication and business acumen; bachelor's degree in CS, Electronics, or CE
Python, SQL, Spark, Kafka, Azure, GCP, AWS
7mo
Save
Mark Applied
Hide
Senior Data Engineer
Riyadh, Riyadh Province, Saudi Arabia
OnsiteFull Time
webook.com
webook.com: Saudi super-app for booking events, travel, and entertainment experiences.
5+ YOE5-6 years as a Data Engineer, design and implement ETL pipelines, data infra (lakes/warehouses/real-time), data modeling and quality, Python/Java, SQL, cloud and DevOps familiarity.
Apache Airflow, Airbyte, AWS, Google Cloud, Elasticsearch, Google BigQuery, MongoDB, Python, Java, SQL, Tableau, Qlik