Citi
Posted 1d ago

Big Data Engineer - Python and Spark

Citi
Pune, Maharashtra, India
HybridFull Time
Responsibilities
  • redesigning processes
  • processing datasets
  • solving problems
Requirements
  • Requires a bachelor's or master's degree
  • Python, Scala, SQL
  • Unix scripting
  • Big data systems
  • Cloud data management, ETL
  • Data warehouses
  • Machine learning, and 0–2 years of experience
  • 3 years of software design experience is also stated
Technical tools mentioned
HiveHadoopSparkPythonScalaSQLUnixTalendPepperdataCloudera

Job description

The Data/Information Mgt Analyst is a trainee professional role. Requires a good knowledge of the range of processes, procedures and systems to be used in carrying out assigned tasks and a basic understanding of the underlying concepts and principles upon which the job is based. Good understanding of how the team interacts with others in accomplishing the objectives of the area. Makes evaluative judgements based on the analysis of factual information. They are expected to resolve problems by identifying and selecting solutions through the application of acquired technical experience and will be guided by precedents. Must be able to exchange information in a concise way as well as be sensitive to audience diversity. Limited but direct impact on the business through the quality of the tasks/services provided. Impact of the job holder is restricted to own job.

Qualifications:

  • Master’s / Engineering Degree with 0- 2 years of experience in Big Data systems, Hive, Hadoop, Spark (Python/ scala) and cloud-based data management technologies
  • Hands-on experience in Unix Scripting, Python and Scala programing along with strong experience in SQL.
  • Comfortable working with completed unstructured, undocumented code and turning it around into best-in-class code redesigning costly compute and data processes and aligning to best development standards
  • Experienced in working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.
  • Well versed with necessary data preprocessing and application engineering skills
  • At least 3 years of experience designing software systems with intense computational needs across real time and batch process .
  • Experience and understanding of Supervised, unsupervised machine learning techniques
  • Exposure to data ingestion, ETL tools such as Talend, modeling tools, Performance Management tooling such as Pepper data, Cloudera stack will be a plus
  • Knowledge of data management, data governance, data security and regulatory practices
  • Ability to identify, clearly articulate and solve complex business problems and present them to the management in a structured and simpler form
  • Should have experience of working in onsite, offsite delivery model
  • Experience working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.
  • Previous related experience preferred
  • High attention to detail


Education:

  • Bachelors/University degree or equivalent experience


This job description provides a high-level review of the types of work performed. Other job-related duties may be assigned as required.

------------------------------------------------------

Job Family Group:

Decision Management

------------------------------------------------------

Job Family:

Data/Information Management

------------------------------------------------------

Time Type:

Full time

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

 

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

About Citi

Providing global banking, investment, and wealth management services.

Similar jobs

Big Data Engineer roles near Pune, Maharashtra
19h
Save
Mark Applied
Hide
Big Data Engineer - Python and Spark
Pune, Maharashtra, India
HybridFull Time
Citi
CitiNYSE: C: A global financial services providing banking and credit services.
0+ YOERequires a bachelor's or master's degree, 0–2 years in big data systems, and experience with Python, Scala, SQL, Hadoop, Hive, Spark, Unix scripting, cloud data management, ETL, and large datasets.
Hive, Hadoop, Spark, Python, Scala, SQL, Unix, Talend, Pepperdata, Cloudera
1d
Save
Mark Applied
Hide
Big Data Engineer - Python and Spark
Pune, Maharashtra, India
HybridFull Time
Citi
CitiNYSE: C: Global diversified financial services holding.
0+ YOERequires a bachelor's or master's/engineering degree, 0–2 years in big data systems, and experience with Python, Scala, SQL, Unix scripting, Spark, Hadoop, Hive, cloud data management, ETL, and machine learning.
Hive, Hadoop, Apache Spark, Python, Scala, SQL, Unix, Talend, Pepperdata, Cloudera
1d
Save
Mark Applied
Hide
Senior Big Data Engineer I
Pune, Maharashtra, India
HybridFull Time
MetLife
MetLifeNYSE: MET: Global provider of insurance, annuities, and financial services.
8+ YOERequires 8–10+ years of relevant experience, a bachelor's degree in computer science, information technology, or equivalent, and expertise in SQL, Python or Scala, big data frameworks, Azure, ETL, and data engineering.
SQL, Python, Scala, HBase, Cosmos DB, Apache Spark, Hadoop, Hive, Azure Data Factory, Event Hubs, Azure Functions, Synapse, Databricks, Git, Azure DevOps, Unix shell scripting, Kafka, MongoDB, NiFi, CI/CD pipelines
1w
Save
Mark Applied
Hide
Big Data Engineer
Pune, Maharashtra, India
OnsiteFull Time
UST
UST: Global provider of digital transformation and IT services.
6+ YOERequires 6–8 years in Big Data or Data Engineering, Java or Scala, Python, Apache Spark, Databricks, AWS, workflow orchestration, scalable pipelines, data architectures, and strong analytical skills.
Scala, Apache Spark, AWS, Amazon EMR, Amazon S3, Python, AWS Glue, Databricks, Java, Apache Airflow, Dagster, Parquet, Avro, ORC, JSON, Git, CI/CD, Claude Code, Docker, Terraform
1w
Save
Mark Applied
Hide
Big Data Hadoop
Pune, Maharashtra, India
RemoteFull Time
Zensar
ZensarNational Stock Exchange of India: ZENSARTECH: Global technology firm providing digital transformation and infrastructure services.
5+ YOERequires 5–12 years of experience with Scala and Apache Spark, functional and object-oriented programming, static typing, and sbt or Maven for dependency management and JAR packaging.
Scala, Apache Spark, sbt, Maven, JAR
2w
Save
Mark Applied
Hide
Senior Big Data Engineer
Pune, Maharashtra, India
OnsiteFull Time
Qualys
QualysNASDAQ: QLYS: Provides cloud-based platform for cybersecurity and compliance management.
5+ YOE5+ years in data engineering, bachelor's in CS/CE or related, experience with Apache Spark and Apache Kafka, scaling high-throughput platforms, and mentoring engineers.
Apache Spark, Apache Kafka, Scala, Java, Python
1mo
Save
Mark Applied
Hide
Big Data Engineer
Bengaluru or Chennai or Hyderabad or Kolkata or Pune or Delhi or Visakhapatnam
OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
5+ YOE5+ years experience with Hadoop ecosystem, Spark, Java/Python/Scala, cloud big-data platforms, and machine learning frameworks; relevant big-data certifications preferred.
Hadoop, HDFS, MapReduce, Hive, Pig, HBase, Spark, Java, Python, Scala, NoSQL, ETL, Flink, Pyspark, AWS EMR, Azure HDInsight, Google Cloud Dataproc, TensorFlow, PyTorch, Scikit-learn
1mo
Save
Mark Applied
Hide
Pyspark Bigdata
Pune, Maharashtra, India
OnsiteFull Time
Virtusa
Virtusa: Global provider of digital engineering and IT outsourcing services.
Experience designing, building and deploying PySpark/Python big data solutions; strong analytical, problem-solving, SDLC knowledge, and communication skills.
PySpark, Python