Amazon
Posted 2w ago

Data Engineer, AOP - RoW Central Data Engineer Team

Amazon
Beijing, Beijing, China
OnsiteFull Time
Responsibilities
  • building pipelines
  • optimizing queries
  • managing incidents
Requirements
  • 1+ years data engineering experience
  • Data modeling and ETL
  • SQL and scripting (Python
  • KornShell)
  • Redshift/Aurora administration
  • AWS services
  • On-call incident response
  • Stakeholder collaboration
Technical tools mentioned
Amazon RedshiftAurora RDSTableau ServerEDXSNSAndesAPIsS3LambdaECS FargateStep FunctionsCloudWatchSQSCDKCloudFormationSecrets ManagerEventBridgeMyUniverseSQLPL/SQLDDLMDXHiveQLSparkSQLScalaPythonKornShellHadoopHiveSparkEMRInformaticaODISSISBODIDatastage

Job description

Description

The RoW Central Data Engineering team builds and operates the central data infrastructure backbone for Amazon's Rest-of-World (ROW) business operations, serving 5,000+ daily users across 14,000+ dashboards and 70,000+ daily data job runs. We run mission-critical systems including a centralized Amazon Redshift cluster, Aurora RDS real-time applications, Tableau Server, and a suite of automated data pipelines that ingest from EDX, SNS, Andes, APIs, and S3.
We are looking for a Professional Data Engineer who takes ownership seriously, thinks clearly under ambiguity, and brings strong technical depth across databases, cloud infrastructure, and data pipeline engineering. You will work alongside senior engineers and technical managers to design, build, and maintain production-grade data systems that directly enable operational decision-making across RoW countries, India, and other emerging markets. If you enjoy untangling complex data problems and taking end-to-end ownership of your solutions, this role is for you.

Key job responsibilities
Design and build data pipelines — architect and implement robust batch, intraday, and near-real-time ETL/ELT pipelines ingesting data from diverse sources including EDX datasets, SNS events, Andes tables, REST APIs, and S3, landing data reliably into Redshift and Aurora.
Own Redshift infrastructure — write, optimize, and tune SQL and ETL workloads on Amazon Redshift; manage WLM queues, distribution/sort keys, materialized views, and query performance; proactively identify and resolve performance bottlenecks.
Manage big data lifecycle — design data models and schemas (star/snowflake), enforce data partitioning and retention policies, implement data quality checks, and ensure data accuracy and freshness SLAs are consistently met.
Deliver on ambiguous requirements — independently break down loosely defined business asks into concrete technical deliverables with clear scope, milestones, and acceptance criteria; drive from requirement to production with minimal hand-holding.
Build automation and self-healing systems — reduce manual toil through automation (Lambda, ECS, Step Functions, CloudWatch alarms); contribute to the team's Server Auto Maintenance Program and ETL cleanup initiatives.
AWS cloud engineering — use AWS services (Redshift, Aurora RDS, S3, Lambda, ECS Fargate, SQS, SNS, CDK/CloudFormation, Secrets Manager, EventBridge) to build scalable, cost-efficient, and maintainable data infrastructure.
Drive cost optimization — proactively identify inefficiencies in SQL workloads, cluster utilization, and pipeline design; propose and implement optimizations that reduce AWS spend without compromising reliability.
Support stakeholders and data consumers — partner with Business Analysts, BI Engineers, Data Scientists, and PMs to understand data needs; deliver clean, documented, raw data pipelines; maintain clear boundaries around pipeline ownership and scope.
Maintain operational excellence — participate in on-call rotation, respond to production incidents with urgency and structured root cause analysis, and implement permanent fixes rather than workarounds.

A day in the life
A typical day for a DE looks like:
Oncall: Review pipeline/infra/services run status on dashboards; triage any failed jobs or data freshness alerts; provide ETA and updates to stakeholders as needed.
Core hours: Work on active sprint deliverables — this may include writing CDK infrastructure code, developing Redshift SQL models, building a new EDX ingestion pipeline, or debugging a WLM contention issue on the central cluster.
Collaboration: Join a sync with stakeholders to understand a new data onboarding request; push back clearly when scope creep or non-standard pipeline patterns are introduced; document the agreed design in the team wiki.
Deep work: Independent heads-down time on complex tasks — performance tuning a slow Redshift query, infra upgrade, refactoring a Lambda trigger handler, or writing a CDK stack for a new ECS data job etc.
Wrap-up: Update task statuses in SIM tickets; code review a peer's PR; document any patterns or learnings into the team's internal knowledge base.
There is no one telling you exactly what to do each hour — you are expected to manage your own task queue, surface blockers early, and keep work moving forward with accountability.

About the team
We are a small, high-impact engineering team established in 2020. Our mission is to provide a unified, highly available, and scalable data infrastructure for Amazon's Rest-of-World operations — covering India, Japan and emerging markets that collectively represent a major and fast-growing segment of Amazon's global business.
We operate at significant scale:
Central Redshift Cluster — 5,000+ daily active users, 70,000+ daily job runs, 14,000+ dashboards on Tableau, QuickSight, and Fusion
Tableau Server — 144 developers, 8,000+ dashboards, 20,000+ daily visits
Aurora RDS — real-time operational applications (capacity alerts, reactive scheduling tools)
GenAI Platform — MyUniverse, our internal data hub, now integrated with Stella 3.0 — an digital AI agent that automates 90%+ of routine operational tasks cross teams. We are a lean team that punches above our weight. Every engineer owns a broad surface area, ships real systems used by thousands of people daily, and is expected to continuously raise the bar — both on the technical side and on how we serve our stakeholders. We value clarity of thought, ownership without ego, and building things that last.

Basic Qualifications

- 1+ years of data engineering experience
- Experience with data modeling, warehousing and building ETL pipelines
- Experience with one or more query language (e.g., SQL, PL/SQL, DDL, MDX, HiveQL, SparkSQL, Scala)
- Experience with one or more scripting language (e.g., Python, KornShell)

Preferred Qualifications

- Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
- Experience with any ETL tool like, Informatica, ODI, SSIS, BODI, Datastage, etc.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

About Amazon

Global online retail and cloud computing technology provider.

Similar jobs

Data Engineer roles near Beijing, Beijing
5d
Save
Mark Applied
Hide
数据研发工程师
Beijing or Hangzhou
OnsiteFull Time
Alibaba
AlibabaNYSE: BABA: Provides online marketplaces, cloud computing, and digital payment services.
Bachelor's degree or higher in computer science, statistics, mathematics, or related fields; big data, data warehouse, distributed architecture, ETL, SQL, UDF, data governance, Python, and LLM or agent application experience.
SQL, Python, LLM
1w
Save
Mark Applied
Hide
数据技术及产品部-RL Data工程师-Work领域(法律方向)
Beijing or Hangzhou
OnsiteFull Time
Alibaba Group
Alibaba GroupNYSE: BABA: A global technology specializing in e-commerce and cloud computing.
3+ YOERequires legal domain expertise, RL/RFT training-data knowledge, taxonomy and rubric design, trajectory annotation and root-cause analysis, Python scripting, and experience producing data at scale.
Python, LLM
1w
Save
Mark Applied
Hide
高级数据工程师(座舱AI)
Beijing, Beijing, China
OnsiteFull Time
Xiaomi
XiaomiHong Kong Stock Exchange: 1810: Designs and manufactures smartphones, consumer electronics, and smart hardware.
3+ YOEBachelor's degree or higher in computer science, software engineering, data science, or related field; 3+ years of data engineering experience; Python, SQL, big data frameworks, ETL, and data modeling expertise.
Large Language Model (LLM), Python, SQL, Spark, Flink, Kafka, Hive, Feature Store, Kubernetes, Airflow, DolphinScheduler
1w
Save
Mark Applied
Hide
数据工程
Shenzhen or Beijing or Shanghai or Guangzhou
OnsiteFull Time
Tencent
TencentHong Kong Stock Exchange: 0700: Multinational technology conglomerate providing internet and entertainment services.
Bachelor's degree or higher in mathematics, statistics, operations research, computer science, or related fields; knowledge of big data technologies, programming, and SQL optimization.
Spark, Flink, TensorFlow, PyTorch, Java, C++, Python, Shell, SQL, GitHub, Large Language Model (LLM)
1w
Save
Mark Applied
Hide
数据工程
Shenzhen or Beijing or Shanghai or Guangzhou
OnsiteFull Time
Tencent
TencentHKEX: 0700: Provides integrated internet services, digital entertainment, and cloud technology.
Bachelor's degree or higher in mathematics, statistics, operations research, computer science, or related fields; knowledge of big data technologies, programming, SQL optimization, and AI tools.
Apache Spark, Apache Flink, TensorFlow, PyTorch, Java, C++, Python, Shell, SQL, GitHub, Large Language Model (LLM), AI
1w
Save
Mark Applied
Hide
海外数据研发工程师
Beijing, Beijing, China
OnsiteFull Time
iReader Technology
iReader TechnologyShanghai Stock Exchange: 603533: Operates a global digital platform for e-books and novels.
Requires big data architecture expertise, Hadoop ecosystem experience, data warehouse and database development, real-time processing with Java or Scala, Linux, Python or Shell, and large-scale data optimization skills.
HDFS, Hadoop, YARN, DataX, Sqoop, HBase, Hive, Spark, StarRocks, Impala, Kylin, MySQL, Redis, ES, BI, DMP, Linux, Python, Shell, Java, Scala, Kafka, Flink, DSP, ADX, SSP, RTB, RTA
2w
Save
Mark Applied
Hide
智能体数据工程师(仿真环境方向)
Beijing, Beijing, China
OnsiteFull Time
ModelBest
ModelBest: Develops efficient, on-device large language models and edge AI.
Proficient Python and engineering practices; experience with containers, remote execution, browser/GUI automation; knowledge of Docker, filesystems, networks, isolation; ability to convert business scenarios into executable, verifiable tasks.
Python, Docker, Kubernetes, Ray
3w
Save
Mark Applied
Hide
大模型数据工程师/专家
Beijing, Beijing, China
OnsiteFull Time
Momenta
MomentaHKEX: Momenta: An AI and autonomous driving building data and infrastructure for large-scale model training and vehicle intelligence.
Design and develop large-scale distributed data platform for multimodal (image, video, lidar) pipelines; proficient in Python/Go/C++; deep experience with TB/PB storage, distributed compute, and large-model training data workflows.
Python, Go, C++