Senior Data Engineer (Databricks / PySpark)

Posted on July 27, 2026

Apply Now

Job Description

Senior Data Engineer (Databricks / PySpark)

Overview

Exp: 9 to 12 Years

Contract for 6- 12 Months+ Extendable.

Location: Bengaluru | Noida | Gurgaon | Mumbai | Pune

Work Location: Remote (India)

Candidates should be flexible to work from an EXL office whenever business requires. While the role is primarily remote, it is not a permanent work-from-home position.

Key Responsibilities

  • Design, build, and productionise batch and incremental ETL/ELT pipelines using Databricks (notebooks, Jobs/Workflows, DLT), PySpark, and Spark SQL across bronze/silver/gold Delta Lake layers.
  • Implement the medallion architecture with Delta Lake features — ACID merges/upserts (SCD1/SCD2), schema evolution/enforcement, time travel, OPTIMIZE/Z-ORDER, vacuuming — and manage data with Unity Catalog.
  • Orchestrate workloads via Databricks Workflows and/or Azure Data Factory; parameterise pipelines, implement restartability, checkpointing, and idempotent re-runs.
  • Tune Spark for performance and cost: partitioning strategy, join/broadcast optimisation, skew handling, caching, cluster sizing/autoscaling, Photon usage, and job-level cost monitoring.
  • Build data quality checks (expectations, reconciliation against source, row/measure-level controls) and implement failure alerting, logging, and lineage-friendly design.
  • Ingest from diverse sources — RDBMS (SQL Server/Oracle), files (CSV/Parquet/JSON/XML), APIs, and streaming/CDC feeds — into ADLS Gen2 landing zones.
  • Contribute to CI/CD for data: Git-based development, pull-request reviews, Azure DevOps/GitHub Actions pipelines, environment promotion, and infrastructure/config as code.
  • Collaborate with Data Modelers, BDAs, testers, and onshore leads in Agile ceremonies; convert mapping specifications into robust, reviewed code; mentor junior engineers.
  • Provide L3 production support for owned pipelines: triage incidents, perform root-cause analysis, and drive permanent fixes within SLAs.

Required Skills

  • 9–12 years overall; strong recent hands-on Databricks and PySpark delivery experience (3+ years Databricks preferred).
  • Expert-level Spark (DataFrame API, Spark SQL, window functions, UDF trade-offs) and advanced Python for data engineering.
  • Advanced SQL — complex transformations, performance tuning, analytical/window queries on large volumes.
  • Delta Lake internals and medallion/lakehouse architecture in production.
  • Azure data stack: ADLS Gen2, Azure Data Factory, Key Vault, and Databricks administration basics (clusters, pools, jobs, secrets).
  • Data warehousing concepts: star schemas, fact/dimension loading patterns, SCD handling, reconciliation.
  • Git-based SDLC with CI/CD exposure (Azure DevOps or GitHub).
  • Insurance or financial-services data experience — ideally P&C (policy/claims/premium/reinsurance structures).

Preferred Skills

  • Delta Live Tables, Unity Catalog governance, Databricks Asset Bundles.
  • Streaming (Structured Streaming, Auto Loader, Event Hubs/Kafka) and CDC tools.
  • dbt on Databricks; Airflow; Terraform.
  • Guidewire, Duck Creek, or London-market source systems; Snowflake or Synapse coexistence.
  • Databricks Data Engineer Associate/Professional certification.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

Other Details

  • Clear written and spoken English; comfortable presenting design decisions to onshore architects and client stakeholders.
  • Proven offshore-delivery discipline: crisp status reporting, proactive risk flagging, dependable overlap-hours availability.
  • Ownership mindset — drives issues to closure without follow-up.

Required Skills

databricks pyspark

Community Discussion

Ask questions, share feedback, or discuss this opportunity. Your comment will be visible to everyone.
Be the first to share your thoughts on this opportunity.

Clarification Board

Your Clarifications
"Send your Job Related Query - you'll get a reply soon."