ZipNom Technologies
Posted 3w ago

Data Engineer

ZipNom Technologies
India
RemoteFull Time
Responsibilities
  • ingesting data
  • parsing data
  • monitoring pipelines
Requirements
  • 3–6 years data engineering/backend experience with production scraping/parsing
  • Strong Python, parser testing, PostgreSQL expertise
  • Operational maturity for nightly pipelines
Technical tools mentioned
PythonCelerylxmlBeautifulSouppandaspdfplumberPostgreSQLAmazon S3Amazon SQSAmazon CloudWatchAmazon SES

Job description




This is a remote position.

About Verizol

Verizol is ZipNom's company intelligence and verification platform: 45+ APIs (KYC, KYB, bank, face, OCR, CKYC, eSign) built on a daily MCA data pipeline, sold through a single API key and prepaid credit wallet. Customers are CA firms, fintechs, NBFCs, and marketplaces. The engineering blueprint is written, the 8-week build plan is set, and the team is small enough that your name is on what ships.

The Role

You own the asset the entire company is built on. Every Verizol product — free search, daily alerts, KYB reports, watchlists — reads from one normalized, indexed database of Indian company registrations, and that database exists because your pipeline ingests, parses, and diffs MCA data every single day. This is the least substitutable seat on the team: when a government data source silently changes format at 2 AM, you're the one whose alarms catch it and whose parsers get fixed before customers notice.

This is not clean-API ETL. It's real-world data engineering against uncooperative sources — scraping, format archaeology, checksum validation, and building a system honest enough to know when its own data is stale.




Requirements

What You Will Do

  • Build source ingestors for MCA data (new incorporations, master data, index of charges, DIN registry, filing metadata) with raw-first storage to S3 — every byte stored before parsing, so parser bugs are always replayable
  • Write parsers as pure, fixture-tested functions: raw bytes in, typed rows out, with a fixture for every format variant ever observed in the wild
  • Build and own the snapshot + diff engine: detect exactly what changed for every company daily, and emit granular change events (director resigned, charge created, status changed) — the event stream that powers Alerts and Watchlists
  • Own the daily production schedule: ingest 02:00 → diff 04:00 → digests 06:30 → alert emails at 08:00 IST, with stage-level retries and a hard rule that stale data never silently ships
  • Build the alerts digest job: filter matching across all subscriber streams, email rendering, SES delivery — for CA subscribers, this email is the product
  • Build the monitoring that keeps you sane: row-count anomaly alarms, parser error rates, per-source freshness gauges
  • Maintain the reference tables (state/ROC mappings, NIC codes) and run the historical backfill (24+ months of incorporations)

Required Skills

  • 3–6 years of data engineering or backend work with production scraping/parsing experience against uncooperative sources — this is the non-negotiable; clean-API ETL alone won't prepare you for MCA
  • Strong Python: Celery (or equivalent task queues), lxml/BeautifulSoup, pandas, pdfplumber or similar document parsing
  • Solid PostgreSQL: bulk upserts, COPY, partitioning, and index design for time-series feed queries
  • Testing instincts for data: fixtures, golden files, replay tests — you can prove a parser change didn't corrupt yesterday
  • Operational maturity: your jobs run while everyone sleeps; you build the alarms first and take the 2 AM page seriously
  • Comfort with proxies, rate budgets, and being a polite, resilient client of fragile infrastructure

Preferred

  • Prior experience with Indian government/registry data: MCA21, GSTN, EPFO, court records, or similar
  • Familiarity with CIN/DIN structures, ROC organization, or corporate filings
  • AWS: S3 lifecycle policies, SQS, CloudWatch metrics/alarms
  • Email deliverability basics (SES, DKIM) — you'll co-own the digest send




About ZipNom Technologies

Provides custom software development and IT consulting services.

Similar jobs

Data Engineer roles
2h
Save
Mark Applied
Hide
Azure Senior Data Engineer
Pune, Maharashtra, India
OnsiteFull Time
HCLTech
HCLTechNational Stock Exchange of India: HCLTECH: Global technology providing digital, engineering, and cloud services.
Requires strong Azure, Databricks, ADF, Synapse, ETL, SQL, and relational database expertise; bachelor's degree in a related field; Python, Spark, client engagement, and Agile experience preferred.
Microsoft Azure, T-SQL, SSIS, SSAS, SSRS, Azure Data Factory (ADF), Azure Databricks, Azure Synapse Analytics, Azure DevOps, Azure Storage, Data Lake, ETL, SQL, Python, Spark APIs, Agile, Scrum
5h
Save
Mark Applied
Hide
Lead Associate - Data / AI
Pune, Maharashtra, India
HybridFull Time
Davies
Davies: Insurance claims management and legal professional services provider.
3+ YOE3+ MgmtBachelor's degree in computer science or related field; 3–5 years managing BI or data teams; experience building data pipelines and architectures with ADF; strong communication and process improvement skills.
Azure Data Factory, Microsoft Fabric, Microsoft SQL Server, Microsoft Power BI
11h
Save
Mark Applied
Hide
Data Engineering Snowflake Lead Engineer _ Vice President _Data Engineering
Bengaluru, Karnataka, India
OnsiteFull Time
Morgan Stanley
Morgan StanleyNYSE: MS: Global financial services firm providing investment and wealth management.
8+ YOERequires 8+ years in data engineering or a related field, with Python, Snowflake Cortex, Databricks, SQL/NoSQL, data modeling, cloud platforms, and data pipeline experience.
Snowflake, Python, SQL/PLSQL, Snowflake Cortex, Apache Spark, Hadoop, SQL, NoSQL, Databricks, AWS, Azure, Kafka, Git, Jupyter, Tableau, Power BI, Collibra
11h
Save
Mark Applied
Hide
Data Engineer Sr. Consultant
Bengaluru, Karnataka, India
OnsiteFull Time
NTT DATA
NTT DATA: Global provider of IT and business consulting services.
Strong Varicent configuration and development, rules and calculations, integrations, SQL, data engineering, APIs, ETL/ELT, transformation, validation, troubleshooting, and communication skills.
Varicent, SQL, APIs, ETL, ELT
11h
Save
Mark Applied
Hide
Lead Engineer - Data
Bangalore or Chennai
OnsiteFull Time
Kyndryl
KyndrylNYSE: KD: Manages and modernizes mission-critical IT infrastructure systems.
10+ YOE10+ years in data engineering and warehousing with Oracle, SQL/PLSQL, ODI, SAP BODS, Python, Dataiku, Unix/Linux, and TWS; experience in banking risk and regulatory reporting, production support, and data operations.
Oracle, SQL/PLSQL, Python, Business Object, Oracle Data Integrator (ODI), TWS, Shell Scripting, Data IKU, Dataiku, SAP BODS, Unix/Linux, AWS, Azure, Power BI, MicroStrategy, Business Objects, Jira, ServiceNow, Git, DevOps, Microsoft, Google, Amazon
11h
Save
Mark Applied
Hide
Senior Data Engineer
Mumbai, Maharashtra, India
OnsiteFull Time
Capgemini
CapgeminiEuronext Paris: CAP: Provides global IT consulting and digital transformation services.
Experienced data engineering professional responsible for leading a team, managing data engineering projects, ensuring technical excellence, and delivering reliable, scalable data solutions.
11h
Save
Mark Applied
Hide
Senior Consultant | Databricks | Mumbai | Engineering
Mumbai or Bangalore
OnsiteFull Time
Deloitte
Deloitte: Professional services firm providing audit, consulting, and advisory services.
5+ YOERequires 5+ years of relevant data engineering experience, Databricks, PySpark, SQL, cloud platform, data architecture, ETL/ELT, DevOps, troubleshooting, and communication skills.
Databricks, PySpark, SQL, Azure, AWS, GCP, Delta Lake, Delta Live Tables (DLT), Unity Catalog, Databricks Workflows, Serverless Compute, GitHub, CI/CD, DevOps, Genie, Mosaic AI, Databricks Data Intelligence Platform, Python
16h
Save
Mark Applied
Hide
Apprentice - Data Engineer
Pune or United States or India or Poland
HybridFull Time, Internship
StoneX Group
StoneX GroupNASDAQ: SNEX: Provides global financial services and market access.
Completing a bachelor's degree in computer science, engineering, mathematics, or related field by July/August 2027; available for a full-time six-month apprenticeship; familiarity with Python, SQL, PySpark, Azure, and Databricks.
Python, SQL, PySpark, Microsoft Azure, Databricks, Power BI, React.js, .NET, SQL Server, Profisee, RedPanda, Kafka, OpenAPI, Airflow, Dynamics, Jenkins, GitHub, Docker, Kubernetes